PokeBot Raises Nine-Figure Pre-A: Embodied AI's Battleground Shifts from Walking to Manipulation
Category: industry · embodied AI / robot manipulation Time: 2026-08-03 08:40 (Beijing Time) Sources: AI Tech Review exclusive; public materials from Tsinghua TEA Lab and Xu Huazhe
Where the Money Goes
According to AI Tech Review, PokeBot, founded in April 2026, has completed a nine-figure (USD 100M-class) Pre-A funding round. Shunwei Capital and Matrix Partners China co-led the round, with participation from Zhongding Capital, JK Capital (Ubiquant), Junshan Capital, SEE Fund, Liepin Investment, and Yuannuo Capital. Existing investors—including Yunqi Capital, Xiaomi Strategic Investment, HLF Fund, Innoangel Fund, and Eastern Gateway Capital—followed on.
The report does not disclose the exact amount, post-money valuation, or closing documents, so the accurate characterization is "nine-figure Pre-A," not any specific dollar figure.
Why It Keeps Emphasizing "Manipulation"
Robot navigation and locomotion already have mature engineering paths, but things get suddenly hard at the kitchen counter: cups slip, tofu crumbles, the state of food in the pan changes, and tools and seasonings must be swapped constantly. One wrong move can cascade into dozens of failed subsequent steps.
PokeBot's publicly stated roadmap is WAM × RL × DATA:
- WAM (World Action Model): predicts how actions will change the environment and helps plan next steps;
- RL (full-pipeline real-robot reinforcement learning): lets the robot keep adjusting based on rewards and failures during real interactions;
- DATA: real-world home-scenario data captured via proprietary collection devices, feeding back into the model and reinforcement learning.
- https://www.163.com/dy/article/L3D5EC2I0511DPVD.html
- http://hxu.rocks/
- https://new.qq.com/rain/a/20260622A0ATAF00?refer=cp_1009
- https://c.m.163.com/news/a/L2GTC5U40511C4AA.html
The most striking public demo is a roughly 9-minute fully autonomous mapo tofu cooking video. The report breaks the difficulty into five dimensions—"long, dexterous, changing, cluttered, precise": long-horizon steps, soft ingredients, real-time state changes, multi-tool coordination, and millimeter-level placement—precisely the parts of home environments that are hardest to standardize.
My Take
The real value of this news is that it pulls embodied AI back from "motions that look right" to "whether the consequences of motions can be controlled." Vision-language-action models let a robot understand "go grab the cup," but that doesn't automatically mean it knows how the cup responds to force, how to recover from a misgrasp, or how to replan mid-task.
That said, a single demo cannot prove generality. When household items, lighting, counter height, or ingredient states change, the hard metrics are success rate, recovery time, and human takeover rate. What PokeBot should publish next is cross-home, cross-object, long-duration operating statistics—not more beautifully edited demos.
Original sources and evidence: