Feedforward foundation models have recently shown remarkable 3D reconstruction capabilities. However, existing models exhibit large tracking drift in long-context streaming reconstruction due to error accumulation. In this paper, we revisit loop closure with streaming reconstruction foundation models to enable accurate, drift-free, kilometer-scale reconstruction. Specifically, our method detects loop candidates through global descriptor retrieval, and constructs loop-conditioned windows to estimate the relative poses between looped frames. Given the observation that our adopted streaming reconstruction backbone produces a globally consistent scale, we optimize all frame poses on the $\mathrm{SE}(3)$ manifold with sequential and loop closure constraints, avoiding the pose graph optimization on the $\mathrm{Sim}(3)$ or higher-dimensional $\mathrm{SL}(4)$ manifolds employed in prior works. Extensive experiments show that our method reduces drift and produces consistent geometry on kilometer-scale sequences, significantly outperforming the state of the art.
Tap the figure to view it at full resolution.
CLoSeR System Overview. Given a streaming monocular sequence, our method processes it in a sliding-window manner. For each window, we first tokenize the frames with a DINO-based encoder and then apply a stack of attention layers, comprising frame attention, sliding-window attention, and TTT layers, to maintain global consistency across the sequence. Upon the arrival of each new window, we additionally extract SALAD descriptors to identify potential loop closures against prior windows. When such candidates are detected, we construct a loop-conditioned window by pairing $N_{\text{win}}/2$ frames from the current window with $N_{\text{win}}/2$ frames from the corresponding prior window, predict the relative poses between loop edges from a single forward pass, and perform pose graph optimization on the $\mathrm{SE}(3)$ manifold to correct tracking drift.
@article{li2026closer,
title = {CLoSeR: Closing the Loop for Long-Context Streaming Reconstruction},
author = {Li, Moyang and Zhu, Zihan and Zhang, Wei and Pollefeys, Marc and Barath, Daniel},
journal = {Advances in Neural Information Processing Systems},
year = {2026}
}