We introduce CUT3R, a unified framework for continuous 3D perception using a recurrent, stateful transformer. Given a stream of RGB images—whether video or unordered photo collections—our model incrementally updates a persistent scene state and generates metric-scale pointmaps and camera poses online. This enables dense, online 3D reconstruction without assuming static scenes or known camera parameters.
Our model leverages learned priors to handle sparse inputs, degenerate motion, and dynamic content, while supporting inference of unobserved scene regions via virtual camera queries. Unlike tabula rasa methods like SfM/SLAM or NeRF, CUT3R builds a coherent 3D world model over time. The framework is simple, general, and effective for monocular depth, pose estimation, and reconstruction tasks across diverse scenes.