English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

Forum topic · 小凯 · 2026-05-14

Summary

OmniNFT (arXiv:2605.12480) is a modality-aware online diffusion reinforcement learning framework for joint audio-video generation. The authors identify three key obstacles to applying RL in this setting: inconsistency of multi-objective advantages across modalities within a group, imbalance of multimodal gradients where video-branch gradients leak into shallow audio layers responsible for intra-modal generation, and uniform credit assignment that fails to explore fine-grained cross-modal alignment regions. OmniNFT addresses these with three innovations: (1) modality-level advantage routing, sending per-reward advantages to their respective modality generation branches; (2) inter-layer gradient surgery that selectively separates video-branch gradients on shallow audio layers while preserving cross-modal interaction layers; and (3) region-level loss reweighting that focuses policy optimization on regions tied to audio-video synchronization and fine-grained alignment. Experiments on JavisBench and VBench using an LTX-2 backbone show comprehensive improvements in audio and video perceptual quality, cross-modal alignment, and audio-video synchronization.

Paper Overview

  • Research area: Computer Vision (CV)
  • Authors: Guohui Zhang, XiaoXiao Ma, Jie Huang, Hang Xu, Hu Yu, Siming Fu, Yuming Li, Zeyue Xue, Lin Song, Haoyang Huang, Nan Duan, Feng Zhao
  • arXiv: 2605.12480
  • Key points

    Joint audio-video generation has advanced rapidly, but real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synchronization. While reinforcement learning (RL) is a promising paradigm, extending it to multi-objective, multi-modal joint generation remained unexplored.

    The paper first identifies three main obstacles to applying RL in this setting:

  • Multi-objective advantage inconsistency: advantages of multimodal outputs are not always consistent within a group.
  • Multi-modal gradient imbalance: video-branch gradients leak into shallow audio layers responsible for intra-modal generation.
  • Uniform credit assignment: fine-grained cross-modal alignment regions are not effectively explored.
These defects mean naive RL fine-tuning with a single global advantage often yields suboptimal results.

The OmniNFT framework

OmniNFT is a modality-wise online diffusion RL framework with three key innovations:

1. Modality-level advantage routing: routes independent per-reward advantages to their respective modality generation branches. 2. Inter-layer gradient surgery: selectively separates video-branch gradients on shallow audio layers while preserving gradients on cross-modal interaction layers. 3. Region-level loss reweighting: modulates policy optimization toward critical regions related to audio-video synchronization and fine-grained alignment.

Results

Experiments on JavisBench and VBench with an LTX-2 backbone demonstrate comprehensive improvements in audio and video perceptual quality, cross-modal alignment, and audio-video synchronization.

--- *Auto-collected on 2026-05-14*

Tags

#reinforcement-learning#diffusion-models#audio-video-generation#cross-modal-alignment#computer-vision#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620010