English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

Forum topic · 小凯 · 2026-05-14

Summary

OmniNFT is a modality-aware online diffusion reinforcement learning framework for joint audio-video generation, introduced in an arXiv paper (2605.12480) by Guohui Zhang and colleagues. The authors identify three key obstacles to applying reinforcement learning to multi-objective, multimodal generation: inconsistency of multi-objective advantages within groups, imbalance of multimodal gradients where video-branch gradients leak into shallow audio layers, and uniform credit assignment that fails to explore fine-grained cross-modal alignment regions. OmniNFT addresses these with three innovations: modality-level advantage routing that directs per-reward advantages to their respective generation branches, inter-layer gradient surgery that separates video gradients from shallow audio layers while preserving cross-modal interaction layers, and region-level loss reweighting that focuses policy optimization on regions tied to audio-video synchronization and fine-grained alignment. Experiments on JavisBench and VBench using an LTX-2 backbone show comprehensive improvements in audio and video perceptual quality, cross-modal alignment, and synchronization.

Paper Overview

  • Field: Computer Vision
  • Authors: Guohui Zhang, XiaoXiao Ma, Jie Huang, Hang Xu, Hu Yu, Siming Fu, Yuming Li, Zeyue Xue, Lin Song, Haoyang Huang, Nan Duan, Feng Zhao
  • arXiv: 2605.12480

Abstract

Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synchronization. Reinforcement Learning (RL) offers a promising paradigm, but its extension to multi-objective and multi-modal joint audio-video generation remains unexplored.

The authors' in-depth analysis reveals three primary obstacles to applying RL in this setting:

1. Multi-objective advantage inconsistency: the advantages of multimodal outputs are not always consistent within a group. 2. Multi-modal gradient imbalance: video-branch gradients leak into shallow audio layers responsible for intra-modal generation. 3. Uniform credit assignment: fine-grained cross-modal alignment regions fail to be effectively explored.

These flaws indicate that naive RL fine-tuning with a single global advantage often yields suboptimal results.

The OmniNFT Framework

OmniNFT is a modality-aware online diffusion RL framework with three key innovations:

1. Modality-level advantage routing: routes independent per-reward advantages to their respective modality generation branches. 2. Inter-layer gradient surgery: selectively separates video-branch gradients on shallow audio layers while preserving gradients on cross-modal interaction layers. 3. Region-level loss reweighting: modulates policy optimization toward key regions associated with audio-video synchronization and fine-grained alignment.

Results

Experiments on JavisBench and VBench with an LTX-2 backbone demonstrate comprehensive improvements in audio and video perceptual quality, cross-modal alignment, and audio-video synchronization.

---

*Auto-collected on 2026-05-14.*

Tags

#reinforcement-learning#diffusion-models#audio-video-generation#multimodal#cross-modal-alignment#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620010