English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation via RL

Forum topic · 小凯 · 2026-04-29

Summary

World-R1 is a reinforcement learning framework that aligns text-to-video generation with 3D constraints without modifying the underlying video model architecture. Instead of injecting 3D priors through costly architectural modifications, World-R1 optimizes video foundation models using Flow-GRPO, leveraging feedback signals from pre-trained 3D foundation models and vision-language models to enforce geometric and structural coherence. The framework introduces a specialized pure-text dataset tailored for world simulation to facilitate this alignment, and adopts a periodic decoupled training strategy that balances rigid geometric consistency with dynamic scene fluidity. Experiments show that the approach significantly improves 3D consistency in generated videos while preserving the base model's original visual quality. Authored by Weijie Wang, Xiaoxuan He, and Youping Gu, the paper was released on arXiv (2504.20698) on April 29, 2025, addressing a key limitation of current video generation models: geometric inconsistencies across frames.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Weijie Wang, Xiaoxuan He, Youping Gu
  • Published: 2025-04-29
  • arXiv: 2504.20698
  • Full Abstract

    Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high computational costs and limit scalability. We propose World-R1, a framework that aligns video generation with 3D constraints through reinforcement learning.

    To facilitate this alignment, we introduce a specialized pure text dataset tailored for world simulation. Utilizing Flow-GRPO, we optimize the model using feedback from pre-trained 3D foundation models and vision-language models to enforce structural coherence without altering the underlying architecture. We further employ a periodic decoupled training strategy to balance rigid geometric consistency with dynamic scene fluidity.

    Key Contributions

  • RL-based alignment: Reinforces 3D constraints in video generation rather than relying on architectural changes, avoiding high computational costs and scalability limitations.
  • World simulation dataset: A specialized pure-text dataset designed to support the alignment of video generation with 3D structure.
  • Flow-GRPO optimization: Uses feedback from pre-trained 3D foundation models and vision-language models as reward signals to enforce structural coherence.
  • Periodic decoupled training: Balances rigid geometric consistency against dynamic scene fluidity during training.

Results

Experiments demonstrate that World-R1 significantly enhances 3D consistency in generated videos while preserving the visual quality of the base video foundation model.

---

*Auto-collected on 2026-04-29.*

Tags

#text-to-video#reinforcement-learning#3d-consistency#flow-grpo#world-models#computer-vision#video-generation#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618874