论文概要
研究领域: CV
作者: Mahesh Reddy, Yashesh Savani, Antoine Mercier
发布时间: 2026-08-17
arXiv: 2508.08545
中文摘要
从退化输入中进行高分辨率图像修复极具挑战性,因为它必须在保持全局结构一致性的同时恢复细粒度的局部细节,尤其是在4K分辨率下,直接基于扩散的修复计算成本高昂且容易产生重复或不一致的纹理。本文提出MagnifiQ,一种渐进式上采样和修复图像的框架,分辨率从1024×1024逐步提升至4096×4096。我们的方法利用预训练的文本到图像扩散模型(如SDXL),通过将原始自注意力层替换为计算成本随图像分辨率线性增长的卷积操作,使其适用于更具可扩展性的高分辨率推理。我们进一步提出渐进式上采样策略,在多个分辨率阶段迭代修复图像,细化每个中间输出,而非直接 hallucinate 最终4K图像,从而改善全局一致性并减少高分辨率伪影。为增强局部细节同时控制内容漂移,MagnifiQ使用块特定的文本提示,在修复过程中提供空间局部化的语义引导。在合成和真实世界退化图像上的大量实验表明,MagnifiQ在感知质量和人类偏好方面优于先前的基于扩散的修复方法,生成更清晰的纹理和更一致的4K结果,同时通过其可扩展的骨干网络和渐进式设计提供了实用的速度-质量权衡。
原文摘要
High-resolution image restoration from degraded inputs is challenging because it must preserve global structural consistency while recovering fine-grained local details, especially at 4K resolution where direct diffusion-based restoration is computationally expensive and prone to repeated or inconsistent textures. In this work, we introduce MagnifiQ, an image restoration framework that progressively upscales and restores images across resolutions, e.g., from 1024x1024 to 4096x4096. Our approach leverages a pre-trained text-to-image diffusion model such as SDXL and adapts it for more scalable high-resolution inference by replacing its original self-attention layers with convolutional operations whose computational cost grows linearly with image resolution. We further propose a progressive u...
自动采集于 2026-08-18
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。