Paper Overview
Field: Computer Vision (CV) Authors: Keyan Hu, Mingtao Wang, Ziyu Zhou, Tiandong Shi, Haifeng Li, Ji Qi, Chao Tao Published: 2026-08-28 arXiv: 2608.28517
Abstract
Cross-modal image translation in remote sensing must preserve source-observed content while matching the target-domain distribution. Existing methods jointly learn the target prior and cross-modal dependence from scarce paired data, overlooking a key asymmetry: only the latter intrinsically requires cross-modal correspondence. We formalize this distinction through conditional-score and denoising-risk analyses and propose Learning the Target Priors Before Image Translation (LTP-BIT), a prior-first paradigm that decouples the two learning tasks. LTP-BIT first learns a target-domain generative prior from large-scale unpaired imagery, then retains the pretrained backbone weights and learns source-conditioned control through P-DART, a parameter-efficient dual-stream architecture. Controlled experiments demonstrate that prior matching and scaling mainly improve target-domain realism, while instance fidelity depends more strongly on conditional adaptation. LTP-BIT achieves state-of-the-art performance on SAR-to-RGB and NIR-to-RGB benchmarks while using only 9.81% task-specific parameters. On QXS-SAROPT, it retains near-full-data instance fidelity with only 25% of the paired samples.
Key Takeaways
- Decoupled learning: Target-domain priors and cross-modal dependence are learned separately, addressing a previously overlooked asymmetry in paired-data usage.
- Theoretical grounding: The distinction is formalized via conditional-score and denoising-risk analyses.
- Parameter efficiency: P-DART enables source-conditioned control while preserving pretrained backbone weights.
- Results: SOTA on SAR-to-RGB and NIR-to-RGB with 9.81% task-specific parameters; strong sample efficiency on QXS-SAROPT (25% paired data).