English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Generative World Renderer: A Large-Scale AAA Game Dataset for Bidirectional Rendering

Forum topic · 小凯 · 2026-04-04

Summary

A paper on arXiv (2504.01263) introduces a large-scale dynamic dataset curated from visually complex AAA games to bridge the domain gap limiting generative inverse and forward rendering in real-world scenarios. Using a novel dual-screen stitched capture method, the authors extracted 4 million continuous frames at 720p/30 FPS of synchronized RGB and five G-buffer channels across diverse scenes, visual effects, and environments, including adverse weather and motion-blur variants. The dataset advances bidirectional rendering by enabling robust in-the-wild geometry and material decomposition and facilitating high-fidelity G-buffer-guided video generation. The authors also propose a novel VLM-based evaluation protocol measuring semantic, spatial, and temporal consistency to assess inverse rendering performance without ground-truth annotations. Experiments show that inverse renderers fine-tuned on this dataset achieve superior cross-dataset generalization and controllable generation, with the VLM evaluation correlating highly with human judgment. A companion toolkit lets users restyle AAA game footage via text prompts using G-buffer editing.

Paper Overview

Research Area: Computer Vision (CV) Authors: Zheng-Hui Huang, Zhixiang Wang, Jiaming Tan Release Date: 2025-04-01 arXiv: 2504.01263

Abstract

Scaling generative inverse and forward rendering to real-world scenarios is bottlenecked by the limited realism and temporal coherence of existing synthetic datasets. To bridge this persistent domain gap, the authors introduce a large-scale, dynamic dataset curated from visually complex AAA games. Using a novel dual-screen stitched capture method, they extracted 4M continuous frames (720p/30 FPS) of synchronized RGB and five G-buffer channels across diverse scenes, visual effects, and environments, including adverse weather and motion-blur variants.

The dataset uniquely advances bidirectional rendering:

  • Inverse rendering: enabling robust in-the-wild geometry and material decomposition
  • Forward rendering: facilitating high-fidelity G-buffer-guided video generation
  • Furthermore, to evaluate the real-world performance of inverse rendering without ground-truth annotations, the authors propose a novel VLM-based evaluation protocol that measures semantic, spatial, and temporal consistency.

    Key Findings

  • Inverse renderers fine-tuned on this dataset achieve superior cross-dataset generalization and controllable generation.
  • The VLM-based evaluation correlates highly with human judgment.
  • Combined with the accompanying toolkit, the forward renderer enables users to edit the style of AAA game footage from G-buffers using text prompts.
  • Links

  • Paper: https://arxiv.org/abs/2504.01263
--- *Auto-collected on 2026-04-04*

Tags

#computer-vision#inverse-rendering#video-generation#datasets#g-buffer#generative-models#aaa-games#vlm-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169521