Paper Overview
Research Area: CV Authors: Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang, et al. Published: 2026-04-03 arXiv: 2604.03225
Summary
Most recent generative image super-resolution (SR) methods rely on adapting large text-to-image (T2I) diffusion models pretrained on web-scale text-image data. While effective, this paradigm starts from a generic T2I generator, despite that SR is fundamentally a low-resolution (LR) input-conditioned image restoration task. In this work, the authors investigate whether an SR model trained purely on visual data can rival T2I-based ones. To this end, they propose VOSR, a Vision-Only generative framework for SR.
Key Ideas
- Semantic features are extracted from the LR input using a pretrained vision encoder, serving as visual semantic guidance that is semantically rich and spatially grounded.
- The authors revisit classifier-free guidance for training generative restoration models and show that the standard unconditional branch is ill-suited to restoration tasks.
- VOSR requires less than one-tenth of the training cost of representative T2I-based SR methods.
- It achieves competitive or better perceptual quality and efficiency in both multi-step and single-step settings.
- On synthetic and real-world benchmarks, it produces more faithful structures with fewer hallucinations.
Results
*Auto-collected on 2026-04-06.*