Key points
On June 30, the topic "robot teachers earning 200+ RMB a day" went viral in Chinese AI circles. It highlights a three-tier talent shortage in China's 2026 embodied intelligence industry:
- Top tier (chief scientists): UBTech offered 124 million RMB/year in 2025, setting a domestic AI executive pay record.
- Middle tier (algorithm/engineering hybrid roles): JD.com opened positions like "data annotation engineer" and "embodied data test development engineer" at 15K–50K RMB/month; a Shenzhen robotics firm offers 12,000 RMB/month for robot data analysts; senior roles reach 300K–350K RMB/year; one "robot after-sales product design expert" role pays 25K–50K RMB/month × 16 months.
- Base tier ("robot teachers"): entry-level teleoperation jobs at 200+ RMB/day (day shift, weekly pay), using joystick-like controllers to guide robots through grasping, opening bottles, folding clothes, etc.
- Two-year goal: 10 million hours of human real-scenario video + 1 million hours of robot body data
- Scenarios: eight categories — home, office, factory, logistics, retail stores, restaurants, healthcare, sanitation
- Scale: 100,000 internal employees + 500,000 external participants
- Positioning: "the largest data collection campaign in human history"
- Job seekers: low-barrier entry (200 RMB/day) plus technical roles at 15K–50K RMB/month.
- Entrepreneurs: data service providers are the scarcest link in the embodied AI value chain.
- Traditional annotation companies: text/image annotation is being automated away, but embodied trajectory data must be human-collected — a structural haven.
- VLA researchers: data scale will accelerate iteration on RT-2 / OpenVLA / Pi0-style architectures.
- Policymakers: a 600,000-person collection effort raises labor, data security, and privacy regulatory questions.
- Data quality: workers trained for under a week; consistency and trajectory cleanliness are concerns — note only ~500K hours of *high-quality* data exist despite broader collection.
- Coverage bias: JD's eight scenarios center on home/office/factory; outdoor, disaster, and extreme-weather long-tail scenarios go largely uncollected.
- Cost: at 500–1,000 RMB/hour, 10 million hours of video implies 5–10 billion RMB.
- Privacy: real-scene video contains faces, documents, home layouts — unclear informed-consent standards and a wide gap from GDPR.
- Hardware bottleneck: 1 million hours of robot body data requires roughly ~340 robots running 24/7 for a year; Mifeng's 2026 target implies thousands of units.
- Open-source leakage: AGIBOT WORLD 2026 could be used directly by overseas model companies — is that a "data export" risk for Chinese firms?
- Baidu Baijiahao: https://baijiahao.baidu.com/s?id=1863614066088886618
- BOSS Zhipin job listing: https://www.zhipin.com/job_detail/9ed98e99870356c903N42t-7E1FZ.html
- Anhui job listing: http://www.ahbys.com/NW/job.html?id=319544
- Tencent News: https://browser.qq.com/mobile/news?doc_id=1286a09157b20452
- CSDN 2026 Q2 embodied AI startup roundup: https://blog.csdn.net/2611_95864581/article/details/160337993
Market price for real-robot data is stable at 500–1,000 RMB/hour (per Yao Maoqing, CEO of Mifeng Technology, an AgiBot incubated company, April 2026).
JD.com's 600,000-person collection campaign
Mifeng Technology targets tens of millions of hours of data capacity in 2026 and tens of billions of hours by 2030.
The data gap
Yao Maoqing's comparison: GPT-5's training corpus ≈ 10 billion hours, while the entire industry's high-quality embodied data ≈ 500,000 hours — a gap of ~10,000x. This gap directly caps embodied model performance.
Meanwhile, AgiBot open-sourced AGIBOT WORLD 2026, positioned as "the world's first real-scenario dataset covering full-domain embodied AI research" — 100% collected from real commercial spaces, hotels, and supermarkets, unlike lab-based predecessors such as DROID and Robomimic.
Analysis: why now
1. VLA model data hunger. Vision-Language-Action models fuse perception, language, and action; lab data can't cover real-world lighting, occlusion, and dynamic environments. 2. Imitation learning is the dominant paradigm. Unlike LLMs that learn passively from scraped text, embodied models must learn from human demonstrations — every trajectory must be produced by a human operator. 3. China's "battle of a hundred robots" phase. April–May 2026 saw massive funding rounds (AgiBot, Galbot, Unitree-affiliated players, Tastone, Lingchu ~2B RMB, Qianxun ~2B RMB, Tastone Intelligence $455M+ Pre-A). First priority after fundraising: buy data.
If JD.com completes its plan, China's embodied data pool could jump from 500K to ~11 million hours by 2027 — shrinking the gap with GPT-5-scale corpora from ~10,000x to ~100x.
Labor-market impact: "robot teacher" could be one of China's largest new job categories in 2026–2027, potentially creating 1–2 million positions industry-wide — a fourth million-scale flexible-employment scenario after food delivery, ride-hailing, and express delivery.