Project Page

Pow3R-SLAM: Real-Time RGB-D SLAM with 3D Reconstruction Priors

Christopher Kolios1, Ishaan Mehta1, Sasa Janjic2, Yeganeh Bahoo1, Sajad Saeedi3

1Toronto Metropolitan University 2University of Windsor 3University College London

Abstract

We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MASt3R-SLAM, a recent work on monocular SLAM using two-view 3D reconstruction priors, we extend the work to incorporate depth as a prior on the network’s prediction, rather than as geometry to fuse. Where traditional RGB-D SLAM systems struggle with sparsity in the depth images, Pow3R utilizes the available depth to give a better-conditioned pointmap, while inferring the depths in empty regions from the two-view photometric, depth, and intrinsic data. Evaluated against MASt3R-SLAM following its protocol on 24 sequences from TUM, 7-Scenes, and Replica, Pow3R-SLAM runs 1.6× faster in wall time, has 15% lower mean trajectory error, a 3.1× lower unscaled error, and produces denser maps, with a 30% lower Chamfer distance. We also introduce a hybrid variant that runs 2.1× faster than MASt3R-SLAM at 25.3 frames per second (FPS), while maintaining improved tracking and mapping accuracy. Against ORB-SLAM3 in RGB-D mode, Pow3R-SLAM is more accurate on TUM, 7-Scenes, and ETH3D-SLAM, and completes every TUM sequence. While Pow3R-SLAM can struggle on a small set of self-similar scenes, its overall performance shows that adding depth as a prior for two-view 3D reconstruction SLAM can be beneficial. Code will be made open-source upon acceptance.

Method Overview

Pow3R-SLAM pipeline diagram. Inputs (intrinsics, keyframe and frame images, sensor depth) feed a front end of depth-conditioned Pow3R, metric scaling and tracking, and a backend of loop closure and depth-anchored global optimization, producing a metric trajectory and a dense metric map. An optional hybrid branch tracks alternate frames by ICP.
Fig. 2. Pow3R-SLAM pipeline. Pow3R-SLAM is an RGB-D SLAM system that uses sensor depth as a prior on a two-view pointmap network. The front end (1-3) runs on every frame. (1) Pow3R predicts pointmaps and confidences for the frame and its keyframe, conditioned on sensor depth and intrinsics (Sec. III-B). (2) Each prediction is made metric against the sensor depth (Eq. 1). (3) The frame is tracked in Sim(3) and may become a keyframe (Secs. III-C–III-D). The backend (4, 5) runs on every new keyframe. (4) Loop closure on Pow3R’s encoder tokens adds edges in familiar locations, and relocalization locates frames the tracker has lost. (5) A global optimization refines every keyframe pose, with keyframe depths anchored to the sensor and confidence-calibrated edges (Sec. III-E.3, Eqs. 5–7). Blue: the sensor-depth path. Orange: optional components, including the hybrid variant, which tracks alternate frames by ICP instead of Pow3R (Sec. III-G), and map-only keyframes (Sec. III-F). ICP frames never become keyframes, so every keyframe comes from a Pow3R pass. Inputs and outputs: TUM fr1/desk.

Primary Demonstration

An overview of the proposed method and main results.

Full system demonstration highlighting core capabilities.

Download the video (MP4, 18.7 MB)

Interactive Map Explorer

Both systems are given the same input frames: MASt3R-SLAM sees the RGB images only, and Pow3R-SLAM sees the same images and the sensor depth. Each map here is gated only by that system’s own confidence, at the percentile printed on its side of the view, with nothing trimmed against the ground truth. Drag the divider to wipe between the two maps. Both halves share one camera, so any pair of points either side of the line is the same viewpoint. Note that some sequences (e.g. fr1/room) include ablation results.

Sensor vs. Network Depth: Holes, Sparsity and Noise

Depth enters Pow3R-SLAM as a prior on the network’s prediction, instead of as geometry to fuse. Thinning the depth to a small fraction of its pixels barely moves the mean trajectory error, because the network fills the gaps from the two views. However, dense but wrong depth drags the prediction with it, and at σ = 0.2 and above ends up worse than supplying no depth at all. Wipe the divider to compare the depth that went in with the depth that came out. With no depth at all the network still returns a complete pointmap, but its metric scale is unconstrained, so the predicted depths land outside the sensor’s own range and saturate the shared colour scale. Pointmap output is from 2 views.

Supplementary Videos

Extended sequences and side-by-side runs.

Paper & Code

Paper

The preprint is being posted to arXiv; the link will appear here.

Code

The code will be released in the Pow3R-SLAM repository after the review process.

Citation

@misc{kolios2026pow3rslam,
  title  = {{Pow3R-SLAM}: Real-Time {RGB-D} {SLAM} with {3D} Reconstruction Priors},
  author = {Kolios, Christopher and Mehta, Ishaan and Janjic, Sasa and
            Bahoo, Yeganeh and Saeedi, Sajad},
  year   = {2026},
  note   = {Preprint}
}

Acknowledgment

We acknowledge the support of the Natural Sciences and Engineering Research Council of Canada (NSERC).