Paper Overview
- Field: Computer Vision (CV)
- Authors: Antoine Guédon, Shu Nakamura, Nicolas Dufour, Jiahui Lei, Ko Nishino, Angjoo Kanazawa
- Published: 2026-06-11
- arXiv: 2606.13644
- Compresses a variable number of unposed RGB views into K latent tokens, forming a single global state.
- Decodes oriented 3D surface points using flow matching.
- An inference-time guidance term injects photometric gradients to correlate nearby points.
- Matches or surpasses feed-forward baselines in reconstruction quality.
- Runs an order of magnitude faster than optimization-based approaches.
- Claimed to be the only feed-forward method combining a global latent representation with arbitrary-resolution decoding.
Abstract
We introduce Surflo, which compresses variable unposed RGB views into K latent tokens (one global state) and decodes oriented 3D surface points via flow matching. An inference-time guidance term correlates nearby points by injecting photometric gradient. Surflo matches or surpasses feed-forward baselines, runs an order of magnitude faster than optimization-based methods, and is the only feed-forward approach combining global latent with arbitrary-resolution decoding.
Key Highlights
*Auto-collected on 2026-06-14*