论文概要
研究领域: CV
作者: Keyan Hu, Mingtao Wang, Ziyu Zhou, Tiandong Shi, Haifeng Li, Ji Qi, Chao Tao
发布时间: 2026-08-28
arXiv: 2608.28517
中文摘要
遥感中的跨模态图像翻译必须保留源观测内容同时匹配目标域分布。现有方法从稀缺配对数据中联合学习目标先验和跨模态依赖性,忽视了一个关键不对称性:只有后者本质上需要跨模态对应。我们通过条件分数和去噪风险分析形式化了这种区别,并提出了LTP-BIT(Learning the Target Priors Before Image Translation),一种先验优先范式,将两个学习任务解耦。LTP-BIT首先从大规模未配对图像中学习目标域生成先验,然后保留预训练骨干权重并通过P-DART学习源条件控制,P-DART是一种参数高效的双流架构。控制实验表明,先验匹配和缩放主要改善目标域真实性,而实例保真度更强烈地依赖于条件适应。LTP-BIT在仅使用9.81%任务特定参数的情况下,在SAR-to-RGB和NIR-to-RGB基准上实现了最先进的性能。在QXS-SAROPT上,它仅使用25%的配对样本就保留了接近全数据的实例保真度。
原文摘要
Cross-modal image translation in remote sensing must preserve source-observed content while matching the target-domain distribution. Existing methods jointly learn the target prior and cross-modal dependence from scarce paired data, overlooking a key asymmetry: only the latter intrinsically requires cross-modal correspondence. We formalize this distinction through conditional-score and denoising-risk analyses and propose Learning the Target Priors Before Image Translation (LTP-BIT), a prior-first paradigm that decouples the two learning tasks. LTP-BIT first learns a target-domain generative prior from large-scale unpaired imagery, then retains the pretrained backbone weights and learns source-conditioned control through P-DART, a parameter-efficient dual-stream architecture. Controlled exp...
自动采集于 2026-09-01
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。