Summary
Lumos-Nexus is a training-efficient unified video generation framework introduced in an arXiv paper (2605.31603) by Jiazheng Xing, Hangjie Yuan, Lingling Cai, Xinyu Liu and colleagues. Connector-based unified video models show strong capabilities in instruction-guided video synthesis, but integrating large high-fidelity generators into a unified training loop is computationally prohibitive, limiting visual quality. Lumos-Nexus addresses this with a two-stage design: during training, only a lightweight generator is aligned with the understanding module; at inference, a Unified Progressive Frequency Bridging (UPFB) module gradually hands generation over to a high-capacity pretrained generator within a shared homogeneous latent space, achieving coarse-to-fine refinement. The paper also introduces VR-Bench, a benchmark for reasoning-driven video generation evaluation. Experiments show Lumos-Nexus achieves significant gains in visual realism and temporal consistency on VBench, along with strong reasoning-driven generation performance on VR-Bench.
Paper Overview
Field: CV / AI
Authors: Jiazheng Xing, Hangjie Yuan, Lingling Cai, Xinyu Liu, et al.
Published: 2026-05-29
arXiv: 2605.31603
PDF: 2605.31603.pdf
Abstract
Connector-based unified video models have demonstrated strong capabilities in instruction-guided video synthesis. However, integrating large high-fidelity generators into a unified training loop is computationally prohibitive, which limits visual quality.
This paper proposes Lumos-Nexus, a training-efficient unified video generation framework that significantly enhances visual fidelity while cultivating reasoning-driven generation capabilities. It adopts a two-stage design:
1. Training time: only the lightweight generator is aligned with the understanding module.
2. Inference time: a Unified Progressive Frequency Bridging (UPFB) module progressively hands generation over to a high-capacity pretrained generator in a shared latent space, enabling coarse-to-fine refinement.
To fill the gap in evaluating reasoning-driven video generation, the authors introduce the VR-Bench benchmark. Experiments show that Lumos-Nexus achieves significant improvements in visual realism and temporal consistency on VBench, while demonstrating strong reasoning-driven generation performance on VR-Bench.
*Auto-collected on 2026-06-02.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177980736