Overview
Research field: Computer Vision (CV) Authors: Ling Xu, Chuyu Han, Borui Li, Hao Wu, Shiqi Jiang, Ting Cao, Chuanyou Li, Sheng Zhong, Shuai Wang Published: 2026-07-02 arXiv: 2607.02501 Categories: cs.RO, cs.CV, cs.OS
Abstract
Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment:
- Multi-rate execution inside closed-loop control
- Latency-first batch-1 inference on heterogeneous hardware
- Extensible embodied interfaces beyond fixed token I/O
- VLA deployments achieved 100.0% and 91.0% task success rates with closed-loop execution
- The WAM benchmark reduced block memory from 312.2 MiB to 88.1 MiB
The Embodied.cpp Runtime
Embodied.cpp is a portable C++ inference runtime for embodied models. Based on an architectural analysis of representative VLA models and WAMs, it captures a shared execution path and organizes it into five layers:
1. Input adapters 2. Sequence builders 3. Backbone execution 4. Head plugins 5. Deployment adapters
The runtime provides modular multi-rate execution, latency-first fused inference, and extensible operator and I/O support, enabling deployment across heterogeneous devices, robots, and simulators through a single backend abstraction.
Results
Evaluations were conducted on two VLA models (HY-VLA and pi0.5) and a preliminary WAM benchmark using LingBot-VA Transformer blocks:
*Source: arXiv:2607.02501*