Marionette: Predicting World States, Rendering Geometry, Painting Appearance
> Paper: Marionette: Predicting World States, Rendering Geometry, Painting Appearance > Authors: Zian Meng, Zhen Li, Chuanhao Li, Qiang Li, Kaipeng Zhang > arXiv: 2608.14530 > Fields: Computer Vision / Game AI / World Models
The Problem: World Models Drift
Current world models typically predict visual observations autoregressively in pixel or latent space. The post illustrates the consequences with a puppet-show metaphor: puppets that sink through the stage floor or drift offstage mirror real failure modes in long-horizon generation.
Key reported numbers from the paper:
- In free-running generation, a hunter–monster pair that stays ~5 m apart in real recordings drifts to 21.2 m apart.
- 33.7% of frames show foot/ground penetration.
- ActionGPT (2.5M params): selects a discrete action token per entity per frame — high-level intent.
- PoseGPT: converts action tokens into continuous joint-level poses.
- Controllability: feeding mismatched action streams on 48 held-out test segments changes root-aligned joint error by 31%; both action-token overrides and direct root scripting work as fine-grained controls.
- State-level repair: adding a terrain collider plus a 6 m separation cap reduces penetration by 66% (33.7% → 11.4%) and brings separation distance to 5.1 m (vs. 21.2 m uncorrected, ~5 m in real data) — with no change to the observation model.
- Fidelity: FVD with recorded poses is 799 vs. 831 with predicted states — no detectable fidelity loss from routing appearance through the predicted state.
- Appearance decay: appearance is an unconstrained quantity with no persistent reference, so it drifts over long horizons; some form of "appearance memory" or style anchor is needed.
- Train/test distribution shift: the observation model sees clean, engine-rendered pose-control videos in training but noisy predicted states at inference.
- Beyond games: extending the explicit-state + deterministic-bridge design to robotics, autonomous driving, and scientific simulation requires domain-specific state definitions and bridges.
The diagnosis: when pose, geometry, and occlusion are all implicitly maintained by a single generative sequence, errors inevitably accumulate in these latent world properties.
Marionette's Three-Layer Architecture
1. Dynamics model (the "skeleton")
A two-stage autoregressive model predicting an explicit, interpretable 276-dim 3D world state (multi-entity joint skeletons, metric root trajectories, 6D root rotations):2. Zero-parameter graphics bridge (the "physical laws")
No learning at all — only fixed, deterministic, differentiable operations:1. Forward kinematics from joint angles and root trajectories 2. Terrain collision detection against a height field 3. Rasterization of 3D skeletons into a pose-control video
> "Nothing here is learned, so nothing here can drift, hallucinate, or need more data."
3. Observation model (the "clothing")
A control-conditioned video diffusion model turns the pose-control video into photorealistic RGB. It handles appearance only — structure is already fixed by the skeleton.Experimental Findings
Limitations and Open Directions
Takeaway
The post argues Marionette represents a philosophical counterpoint to end-to-end learning: some structure should be given, not learned. Geometry is a physical necessity that needs no data; only appearance and action semantics truly need to be learned. The payoff is interpretability — every component emits human-readable objects (action tokens, joint angles, 3D geometry, pixels), so failures can be localized layer by layer.
> "Marionette separates the parts that must be exact from the part that must look right."
References
1. Meng, Z., Li, Z., Li, C., Li, Q., & Zhang, K. (2026). *Marionette: Predicting World States, Rendering Geometry, Painting Appearance*. arXiv:2608.14530. 2. Ha, D., & Schmidhuber, J. (2018). World models. NeurIPS 2018. 3. Hafner, D., et al. (2019). Dream to control: Learning behaviors from latent imagination. 4. Yan, W., et al. (2023). VideoGPT: Video generation using VQ-VAE and transformers. 5. Ho, J., et al. (2022). Imagen video: High definition video generation with diffusion models. 6. Guo, C., et al. (2023). Animating human images with appearance-aware diffusion.