At the 2026 World Robot Conference (WRC) main forum in Beijing, Wang He, founder and CTO of Galaxy General, presented a concrete roadmap for when embodied AI reaches its "ChatGPT moment" — and how the company plans to get there by 2028.
Key points
- Definition of the "ChatGPT moment": robots achieving 70-80% zero-shot success on common skills they were never specifically trained on, with ordinary users able to teach new behaviors without an algorithm background.
- Timeline: Wang He judges current embodied foundation models to be at the "ChatGPT-2 stage," and claims Galaxy General will reach the embodied equivalent of GPT-3.5 to GPT-4 by 2028.
- Commercial target: tens of thousands of humanoid robots deployed across industries by the end of the 15th Five-Year Plan period.
- VLA (Vision-Language-Action) models absorb only action-labeled data; world models absorb massive unlabeled video but cannot directly output joint commands.
- Galaxy General proposes WAM (World-Action Model), which trains on both action-labeled robot data and internet-scale unlabeled human video under a unified objective.
- Milestones: first WAM paper (March 2025) accepted to ICCV 2025; 171 follow-up papers citing the paradigm; NVIDIA's head of embodied AI publicly called WAM "the lifelong game of robotics" in a Sequoia interview.
- AstraBrain WAM 0.5 models the future in latent space — imagining outcomes stripped of lighting and task-irrelevant textures — and supports cross-scene, cross-embodiment, multi-task operation.
- Trained on 2 billion frames of human motion-capture data; Wang He reports a verified scaling law: loss keeps decreasing with data volume at the 100,000-hour scale.
- Demonstrated capabilities: high-dynamic dance (handstands, human-robot duets at the Galbot ET1 launch), teleoperated daily tasks (feeding cats, scooping litter), emergency response, industrial assistance, and retail grasping.
- Compared against NVIDIA's whole-body control: lower latency, more precise motions, better stability on difficult balances.
- L2 human data: 1 million hours total — 500,000 hours of bare-hand manipulation plus 500,000 hours via the new Astra Set head-mounted dual-hand UMI capture rig.
- L5 real-robot return-flow data: 80,000 hours from deployments in healthcare, CATL production lines, and pharmacies.
- Target: scale data volume 10x by 2028 to approach the ChatGPT-moment threshold.
- Robots have picked over 1 million medication boxes across nearly 100 pharmacies in Beijing, Shanghai, Guangzhou, Shenzhen and other cities since December 2024.
- Mechanical fault-free operation has reached 18 months and counting.
- On August 9, Beijing's first provincial/ministerial-level key lab on embodied intelligence foundation models was unveiled in Haidian, co-built by Galaxy General, Peking University, and CATL, with Wang He as director.
- Wang Xingxing (Unitree): the tipping point is a robot completing 80% of tasks in any unfamiliar environment via voice — "two to three years at best, five to ten at worst."
- Huang Qingqiu (Moqi Intelligence): data is the industry's bottom-line constraint — a chicken-and-egg problem.
- Wang He (Galaxy General): 2028 for the embodied ChatGPT moment.
The WAM paradigm
AstraBrain-WBC 0.5: the general "cerebellum"
One brain, three bodies
| Robot | Type | Use case | | --- | --- | --- | | Galbot ET1 | Bipedal humanoid | Real-time multimodal interaction, high-dynamic dance | | Galbot G1 | Wheeled semi-humanoid | Breakfast making, flexible folding, re-planning under interference | | Galbot S1 | Heavy-load wheeled | CATL production line, 7×24h industrial tasks |
At the WRC competitions, the same AstraBrain stack won gold in all three flagship categories — home, retail, and restaurant — including a 6-minute-55-second autonomous restaurant run with zero failures.
Data infrastructure: the five-level pyramid
Commercial traction
WAM-TTT: post-training from customer video
WAM-TTT (Test-Time Training) slashes adaptation cost: a customer wearing a head-mounted camera records ~10 minutes of work video; the model is then fine-tuned with no labels, no teleoperation, and no action data — yielding a scenario-specialized model from the same base brain. This is Galaxy General's answer to the "last mile" of industry-wide deployment.
Two-step definition of the goal
1. Step 1 (2028): 70-80% zero-shot generalization on untrained common skills. 2. Step 2 (by end of the 15th Five-Year Plan): simple post-training lifting any customer scenario to near-100% task success.
Diverging industry timelines at WRC 2026
Observations
The talk did three concrete things: pushed WAM from paper to production (driving three embodiments), turned the "ChatGPT moment" from a slogan into a measurable metric (70-80% zero-shot + 100% post-trained), and hard-wired a 2028 deadline backed by verifiable data assets. Open questions: whether a GPT-2-to-GPT-3.5-style leap in just two years is realistic (it took four years in LLMs), how sharply WAM's boundaries versus VLA and world models can be independently validated, and how WAM-TTT will handle variable customer data quality.
Worth tracking over the next 6-12 months: WAM adoption among leading embodied AI companies, the release cadence toward AstraBrain WAM 1.0, and the real-world robustness behind the WRC competition results.