English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Embodied.cpp: A Portable C++ Inference Runtime for Embodied AI Models on Heterogeneous Robots

Forum topic · 小凯 · 2026-07-06

Summary

Embodied.cpp (arXiv:2607.02501) is a portable C++ inference runtime designed for deploying embodied AI models, including vision-language-action (VLA) models and world-action models (WAMs), on heterogeneous robots and edge devices. Existing inference runtimes are built for request-response serving and fail to meet embodied deployment requirements such as multi-rate execution within closed-loop control, latency-first batch-1 inference on heterogeneous hardware, and extensible I/O interfaces. Embodied.cpp addresses this by extracting a shared execution path from representative VLA models and WAMs and organizing it into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. It provides modular multi-rate execution, latency-first fused inference, and extensible operators and I/O via a single backend abstraction across devices, robots, and simulators. Evaluations on two VLA models (HY-VLA and pi0.5) and an initial WAM benchmark using LingBot-VA Transformer blocks achieved 100.0% and 91.0% task success rates, and reduced per-block memory from 312.2 MiB to 88.1 MiB respectively.

Overview

Research field: Computer Vision (CV) Authors: Ling Xu, Chuyu Han, Borui Li, Hao Wu, Shiqi Jiang, Ting Cao, Chuanyou Li, Sheng Zhong, Shuai Wang Published: 2026-07-02 arXiv: 2607.02501 Categories: cs.RO, cs.CV, cs.OS

Abstract

Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment:

  • Multi-rate execution inside closed-loop control
  • Latency-first batch-1 inference on heterogeneous hardware
  • Extensible embodied interfaces beyond fixed token I/O
  • The Embodied.cpp Runtime

    Embodied.cpp is a portable C++ inference runtime for embodied models. Based on an architectural analysis of representative VLA models and WAMs, it captures a shared execution path and organizes it into five layers:

    1. Input adapters 2. Sequence builders 3. Backbone execution 4. Head plugins 5. Deployment adapters

    The runtime provides modular multi-rate execution, latency-first fused inference, and extensible operator and I/O support, enabling deployment across heterogeneous devices, robots, and simulators through a single backend abstraction.

    Results

    Evaluations were conducted on two VLA models (HY-VLA and pi0.5) and a preliminary WAM benchmark using LingBot-VA Transformer blocks:

  • VLA deployments achieved 100.0% and 91.0% task success rates with closed-loop execution
  • The WAM benchmark reduced block memory from 312.2 MiB to 88.1 MiB
These results show that Embodied.cpp improves deployment efficiency while maintaining high accuracy, and is applicable across diverse embodied model architectures.

*Source: arXiv:2607.02501*

Tags

#embodied-ai#inference-runtime#vla#robotics#cpp#edge-computing#world-action-models#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178209076