English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference (AMNet)

Forum topic · 小凯 · 2026-06-11

Summary

AnyMod-LLVE introduces AMNet, a unified multimodal framework for low-light video enhancement (LLVE) that supports flexible, modality-agnostic inference. Existing multimodal LLVE methods improve enhancement quality by leveraging auxiliary modalities such as event streams and infrared images, but they typically assume these modalities are available at inference time — an assumption rarely valid in real-world deployments. AMNet addresses this with a Spatial-Spectral Dual-Gated Translator that learns the correspondence between auxiliary modalities and RGB inputs, generating implicit auxiliary representations that enable robust enhancement even when auxiliary signals are missing. Trained on RGB-only data combined with large-scale multimodal pretraining, the model handles arbitrary combinations of input modalities at inference and achieves strong performance under missing-modality conditions. The paper (arXiv:2606.11186) is authored by Hangfeng Liang, Yutao Hu, Yanhan Hu, Xiaohan Wu, Wenqi Shao, and Ying Fu.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Hangfeng Liang, Yutao Hu, Yanhan Hu, Xiaohan Wu, Wenqi Shao, Ying Fu
  • Published: 2026-06-09
  • arXiv: 2606.11186
  • Abstract (translated summary)

    AnyMod-LLVE proposes the AMNet framework for low-light video enhancement (LLVE), supporting flexible, modality-agnostic inference. A Spatial-Spectral Dual-Gated Translator learns the correspondence between auxiliary modalities and RGB inputs, generating implicit auxiliary representations for robust enhancement. Built on RGB-only datasets and large-scale multimodal pretraining, the model can handle arbitrary modality combinations at inference time and performs well when auxiliary modalities are missing.

    Original Abstract (excerpt)

    Low-light video enhancement (LLVE) remains a challenging task due to severe information degradation under low-illumination conditions. Recent multimodal approaches have significantly improved enhancement performance by incorporating auxiliary modalities, such as event streams and infrared images. However, these methods typically assume the availability of these modalities at inference, which is often not feasible in real-world scenarios. To solve this problem, in this work, we propose AMNet, a unified multimodal framework for LLVE, to support flexible modality-agnostic inference, where auxiliary modalities may be unavailable. To address the issue of modality absence, we introduce a Spatial-Spectral Dual-Gated Translator that learns the correspondence between auxiliary modalities and RGB in...

    Key Contributions

  • Modality-agnostic inference: AMNet works with any combination of available input modalities, removing the strict requirement of auxiliary signals (e.g., events, infrared) at test time.
  • Spatial-Spectral Dual-Gated Translator: learns cross-modal correspondence between auxiliary modalities and RGB, producing implicit auxiliary representations when real auxiliary data is absent.
  • Training strategy: combines RGB-only data with large-scale multimodal pretraining, yielding robust performance under missing-modality conditions.
---

*Auto-collected on 2026-06-11.*

Tags

#low-light-video-enhancement#computer-vision#multimodal-learning#arxiv#paper#video-enhancement#event-camera#infrared-imaging

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981075