Primary Demonstration
An overview of the proposed method and main results.
Full system demonstration highlighting core capabilities.
Project Page
1Toronto Metropolitan University 2University of Windsor 3University College London
We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MASt3R-SLAM, a recent work on monocular SLAM using two-view 3D reconstruction priors, we extend the work to incorporate depth as a prior on the network’s prediction, rather than as geometry to fuse. Where traditional RGB-D SLAM systems struggle with sparsity in the depth images, Pow3R utilizes the available depth to give a better-conditioned pointmap, while inferring the depths in empty regions from the two-view photometric, depth, and intrinsic data. Evaluated against MASt3R-SLAM following its protocol on 24 sequences from TUM, 7-Scenes, and Replica, Pow3R-SLAM runs 1.6× faster in wall time, has 15% lower mean trajectory error, a 3.1× lower unscaled error, and produces denser maps, with a 30% lower Chamfer distance. We also introduce a hybrid variant that runs 2.1× faster than MASt3R-SLAM at 25.3 frames per second (FPS), while maintaining improved tracking and mapping accuracy. Against ORB-SLAM3 in RGB-D mode, Pow3R-SLAM is more accurate on TUM, 7-Scenes, and ETH3D-SLAM, and completes every TUM sequence. While Pow3R-SLAM can struggle on a small set of self-similar scenes, its overall performance shows that adding depth as a prior for two-view 3D reconstruction SLAM can be beneficial. Code will be made open-source upon acceptance.
An overview of the proposed method and main results.
Full system demonstration highlighting core capabilities.
Both systems are given the same input frames: MASt3R-SLAM sees the RGB images only, and Pow3R-SLAM sees the same images and the sensor depth. Each map here is gated only by that system’s own confidence, at the percentile printed on its side of the view, with nothing trimmed against the ground truth. Drag the divider to wipe between the two maps. Both halves share one camera, so any pair of points either side of the line is the same viewpoint. Note that some sequences (e.g. fr1/room) include ablation results.
Depth enters Pow3R-SLAM as a prior on the network’s prediction, instead of as geometry to fuse. Thinning the depth to a small fraction of its pixels barely moves the mean trajectory error, because the network fills the gaps from the two views. However, dense but wrong depth drags the prediction with it, and at σ = 0.2 and above ends up worse than supplying no depth at all. Wipe the divider to compare the depth that went in with the depth that came out. With no depth at all the network still returns a complete pointmap, but its metric scale is unconstrained, so the predicted depths land outside the sensor’s own range and saturate the shared colour scale. Pointmap output is from 2 views.
Extended sequences and side-by-side runs.
The preprint is being posted to arXiv; the link will appear here.
The code will be released in the Pow3R-SLAM repository after the review process.
@misc{kolios2026pow3rslam,
title = {{Pow3R-SLAM}: Real-Time {RGB-D} {SLAM} with {3D} Reconstruction Priors},
author = {Kolios, Christopher and Mehta, Ishaan and Janjic, Sasa and
Bahoo, Yeganeh and Saeedi, Sajad},
year = {2026},
note = {Preprint}
}
We acknowledge the support of the Natural Sciences and Engineering Research Council of Canada (NSERC).