Overview
- Field: Computer Vision (CV)
- Authors: Ling Xu, Chuyu Han, Borui Li
- Published: 2026-07-04
- arXiv: 2507.03242
- VLA deployments achieved task success rates of 100.0% (HY-VLA) and 91.0% (pi0.5)
- The WAM benchmark reduced block memory from 312.2 MiB to 88.1 MiB
Abstract (translated)
Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment: multi-rate execution inside closed-loop control, latency-first batch-1 inference on heterogeneous hardware, and extensible embodied interfaces beyond fixed token I/O.
The authors present Embodied.cpp, a portable C++ inference runtime for embodied models. Based on an architectural analysis of representative VLA models and WAMs, Embodied.cpp captures a shared execution path and organizes it into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. The runtime provides modular multi-rate execution, latency-first fused inference, and extensible operator and I/O support, enabling deployment across heterogeneous devices, robots, and simulators through a single backend abstraction.
Results
Evaluation on two VLA models (HY-VLA and pi0.5) and a preliminary WAM benchmark based on LingBot-VA Transformer blocks shows:
Original Abstract
Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment: multi-rate execution inside closed-loop control, latency-first batch-1 inference on heterogeneous hardware, and extensible embodied interfaces beyond fixed token I/O.
---
*Auto-collected on 2026-07-05.*