English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VOSR: A Vision-Only Generative Model for Image Super-Resolution

Forum topic · 小凯 · 2026-04-06

Summary

VOSR is a vision-only generative framework for image super-resolution (SR) proposed by Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang and colleagues (arXiv:2604.03225). Unlike most recent generative SR methods that adapt large text-to-image (T2I) diffusion models pretrained on web-scale text-image data, VOSR investigates whether an SR model trained purely on visual data can rival T2I-based approaches. The framework extracts semantically rich, spatially grounded features from the low-resolution input using a pretrained vision encoder as visual semantic guidance, and revisits classifier-free guidance for training generative restoration models, showing that the standard unconditional branch is ill-suited to restoration tasks. VOSR requires less than one-tenth of the training cost of representative T2I-based SR methods while achieving competitive or better perceptual quality in both multi-step and single-step settings, and produces more faithful structures with fewer hallucinations on synthetic and real-world benchmarks.

Paper Overview

Research Area: CV Authors: Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang, et al. Published: 2026-04-03 arXiv: 2604.03225

Summary

Most recent generative image super-resolution (SR) methods rely on adapting large text-to-image (T2I) diffusion models pretrained on web-scale text-image data. While effective, this paradigm starts from a generic T2I generator, despite that SR is fundamentally a low-resolution (LR) input-conditioned image restoration task. In this work, the authors investigate whether an SR model trained purely on visual data can rival T2I-based ones. To this end, they propose VOSR, a Vision-Only generative framework for SR.

Key Ideas

  • Semantic features are extracted from the LR input using a pretrained vision encoder, serving as visual semantic guidance that is semantically rich and spatially grounded.
  • The authors revisit classifier-free guidance for training generative restoration models and show that the standard unconditional branch is ill-suited to restoration tasks.
  • Results

  • VOSR requires less than one-tenth of the training cost of representative T2I-based SR methods.
  • It achieves competitive or better perceptual quality and efficiency in both multi-step and single-step settings.
  • On synthetic and real-world benchmarks, it produces more faithful structures with fewer hallucinations.
---

*Auto-collected on 2026-04-06.*

Tags

#super-resolution#generative-models#computer-vision#diffusion-models#vision-only#image-restoration#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169576