PhiZero: Reason in 'Physical Language' First, Then Render Video
On July 30, a team from the Institute of Automation, Chinese Academy of Sciences (NLPR/CASIA) released the PhiZero preprint (arXiv:2607.28624, project page: phi-zero.github.io), offering a new paradigm for physical AI, world models, and embodied intelligence: reason about the future in a "Physical Language" first, then render the reasoning into video, rather than predicting pixels directly.
PhiZero achieved SOTA on four physics benchmarks:
- Physics-IQ Verified: IQ-Score 41.2 (ahead of Cosmos3-Super 39.5, Grok-Video 34.8, Sora 2 26.5)
- PhyGround: Physics Score 3.01 (ahead of Veo3.1 2.85, Wan2.2-14B 2.90)
- WorldModelBench: Physics Adherence 4.88 (ahead of Wan2.2-5B 4.51, Runway 4.27)
- IntPhys2: Overall 56.34 (ahead of Gemini-2.5 Flash 55.63, GPT-4o 53.75)
- Visual fidelity keeps improving, but physical consistency stagnates
- Performance is excellent for short horizons (1–2 seconds), but collapses beyond 5 seconds
- Predictions of "hidden states" (post-collision bounce, gravity-induced deformation, fluid motion) are barely learned in pixel space
- Visual details in training data saturate the representation capacity, drowning out the "world state transitions" the model should actually learn
- Input: two adjacent video frames
- Processing: encode to latent with the Wan2.2 VAE; extract "transition features between adjacent latent states" with a transition-level Q-Former
- Discretization: compress transition features into discrete symbols via FSQ (Finite Scalar Quantization) — a 25K-entry atomic vocabulary
- Decoding: a Wan2.2-5B diffusion decoder renders the discrete symbols plus the first frame back into video
- Key trick: the diffusion decoder's pretrained prior recovers "visual details" (texture, lighting), while the first frame anchors "scene appearance" — so the discrete bottleneck can focus on learning "state transitions" (motion, collision, deformation)
- Initialization: Qwen3-VL-4B (vision-language model)
- Input: first frame + text action intent
- Processing: autoregressively predicts future steps of the "physical language" sequence
- Output: the sequence is passed to the diffusion decoder to render video
- Initial pool: 50K hours of in-the-wild real video (YouTube / public datasets)
- Progressive filtering → 10K hours for tokenizer pretraining
- Real + simulated video → 5 million 4-second clips for tokenizer SFT and reasoner pretraining
- Further filtering → 1 million high-quality, motion-rich, physics-informative clips for reasoner SFT
- Before LLVM, compilers generated assembly directly: non-portable, hard to debug
- After LLVM, an intermediate representation (IR) decoupled front ends (language → IR) and back ends (IR → hardware), collapsing migration costs
- Before PhiZero, world models were end-to-end black boxes (video → video)
- After PhiZero, there is an intermediate representation (Physical Language), with the reasoner (language → physical language) and renderer (physical language → video) decoupled
- New physical language tokenizers (different quantization algorithms / visual encoders)
- New physical language reasoners (different VLM initializations, different instruction formats)
- New physical language renderers (different video generation backends, 4D Gaussian Splatting, NeRF, 3D)
- Shared physical language vocabularies (a standard library of physical symbols)
- World model teams: evaluate whether PhiZero's tokenizer plugs into your video generation backbone. The paper validates the Wan2.2 + Qwen3-VL path; you can likely reuse components.
- VLA / embodied decision-making teams: physical language is a very lightweight intermediate representation that strips visual redundancy and keeps only "state transitions." This may be more cost-effective than simply training larger VLAs.
- Simulation / Sim2Real teams: the reason-then-render split lets controllability of simulated scenes and realism of visual rendering be optimized independently for the first time.
- AI for Science teams: discrete symbolic representations of physics align naturally with symbolic systems for atoms, molecules, and chemical reactions — a potential "world model + scientific discovery" crossover.
- Model evaluators: physical consistency and visual fidelity are independent optimization dimensions. A visually beautiful model does not necessarily understand physics. Add specialized benchmarks like Physics-IQ, PhyGround, and WorldModelBench.
- PhiZero project page (demos / videos): https://phi-zero.github.io/
- arXiv:2607.28624 HTML version: https://arxiv.org/html/2607.28624
- arXiv TLDR summary: https://arxivtldr.org/abs/2607.28624
- AI Base coverage: https://www.aibase.com/news/30037
- The NextGen Tech Insider coverage: https://www.thenextgentechinsider.com/pulse/phizero-introduces-physical-language-for-advanced-world-modeling
- Full July 30 AI digest: https://www.cnblogs.com/vibe234/p/22114799
This is the first time a world model beats the full circle of baselines — Sora 2, Veo 3.1, Wan2.2, Cosmos-3 — on the "physical plausibility" dimension, while PhiZero's backbone uses far fewer resources.
The Problem: Direct Pixel Prediction Has Hit Its Ceiling
Over the past two years, the mainstream approach in the physical AI / world model space has been "directly predicting future video pixels" — VideoPoet, Genie, Sora, Wan2.2, Cosmos, and Veo are all variants. This paradigm's ceiling became obvious in 2026 H1:
PhiZero's angle: turn the human brain's habit of abstracting how the world works into model architecture. The paper opens with Sherlock Holmes' classic line — "You see, but you do not observe." The brain does not first learn visual details; it abstracts "predictive structure" from visual experience and uses language to organize that structure for explicit reasoning. LLMs have proven that symbolic reasoning spaces are scalable in the digital domain, but natural language is too coarse for physical state transitions.
So PhiZero proposes an intermediate layer — Physical Language: a compact, discrete representation of "world state transitions," learned self-supervised from unlabeled in-the-wild video.
Architecture: The Two-Stage Reason-Then-Render Design
PhiZero consists of two modules:
Module 1: Physical Language Tokenizer
Module 2: Physical Language Reasoner
The paradigm is named reason-then-render — first reason out "future world evolution" as discrete symbols in physical language space, then hand the symbols to the diffusion decoder for rendering. This is fundamentally different from mainstream pixel-space direct prediction: dynamics and rendering are decoupled.
Training Data: From 50K Hours of Raw Video to 5M 4-Second Clips
This "broad pretraining + curated fine-tuning" two-stage recipe follows mainstream NLP practice — a sign that world model training is maturing methodologically.
Why It Matters: The 'Compiler Mindset' Arrives for World Models
PhiZero's real contribution is shifting world models from the "direct pixel prediction" paradigm to a "compiler" paradigm — the same story as the last decade of programming languages:
Within the next 12 months, expect an engineering ecosystem to form around physical language:
Once this ecosystem forms, world models become modular products — just as NLP split into tokenizers, transformer blocks, task heads, and prompt formats after BERT / GPT.
Practical Advice for Embodied AI / World Model Teams
One-Sentence Summary
PhiZero moves world models from "directly predicting pixels" to "reasoning in physical language first, then rendering pixels" — just as programming moved from direct assembly to IR-first compilation. The significance is not any single benchmark win, but the first concretization, standardization, and engineering-ization of an intermediate representation for world models. In 2026 H2, ecosystem-building around physical language may matter more industrially than any single model's score gains.
---
References