CLoSeR

Closing the Loop for Long-Context Streaming Reconstruction

NeurIPS 2026

1ETH Zürich 2University of Stuttgart 3Microsoft
*Equal contribution. Author order is interchangeable.
CLoSeR teaser on KITTI 00

Tap the figure to view it at full resolution.

On KITTI 00 (3.7 km / 4542 frames, several loops), the submap-based approach VGGT-SLAM 2.0 exhibits erroneous point cloud alignment (ATE 108.39 m), while the feedforward method LoGeR suffers from tracking drift due to the absence of explicit constraints when revisiting regions (ATE 62.34 m). In contrast, CLoSeR closes the loops and consistently produces accurate reconstructions at kilometer scale (ATE 6.51 m).

TL;DR CLoSeR brings loop closure to streaming reconstruction: novel loop-conditioned windows let the backbone itself estimate loop constraints while preserving local context and scale, and a lightweight $\mathrm{SE}(3)$ pose graph removes drift at kilometer scale.

Abstract

Feedforward foundation models have recently shown remarkable 3D reconstruction capabilities. However, existing models exhibit large tracking drift in long-context streaming reconstruction due to error accumulation. In this paper, we revisit loop closure with streaming reconstruction foundation models to enable accurate, drift-free, kilometer-scale reconstruction. Specifically, our method detects loop candidates through global descriptor retrieval, and constructs loop-conditioned windows to estimate the relative poses between looped frames. Given the observation that our adopted streaming reconstruction backbone produces a globally consistent scale, we optimize all frame poses on the $\mathrm{SE}(3)$ manifold with sequential and loop closure constraints, avoiding the pose graph optimization on the $\mathrm{Sim}(3)$ or higher-dimensional $\mathrm{SL}(4)$ manifolds employed in prior works. Extensive experiments show that our method reduces drift and produces consistent geometry on kilometer-scale sequences, significantly outperforming the state of the art.

Method

CLoSeR system overview

Tap the figure to view it at full resolution.

CLoSeR System Overview. Given a streaming monocular sequence, our method processes it in a sliding-window manner. For each window, we first tokenize the frames with a DINO-based encoder and then apply a stack of attention layers, comprising frame attention, sliding-window attention, and TTT layers, to maintain global consistency across the sequence. Upon the arrival of each new window, we additionally extract SALAD descriptors to identify potential loop closures against prior windows. When such candidates are detected, we construct a loop-conditioned window by pairing $N_{\text{win}}/2$ frames from the current window with $N_{\text{win}}/2$ frames from the corresponding prior window, predict the relative poses between loop edges from a single forward pass, and perform pose graph optimization on the $\mathrm{SE}(3)$ manifold to correct tracking drift.

Results

BibTeX

@article{li2026closer,
  title   = {CLoSeR: Closing the Loop for Long-Context Streaming Reconstruction},
  author  = {Li, Moyang and Zhu, Zihan and Zhang, Wei and Pollefeys, Marc and Barath, Daniel},
  journal = {Advances in Neural Information Processing Systems},
  year    = {2026}
}