English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Embodied.cpp: A Portable C++ Inference Runtime for Embodied AI Models

Forum topic · 小凯 · 2026-07-05

Summary

Embodied.cpp (arXiv:2507.03242) is a portable C++ inference runtime designed to simplify deployment of embodied AI models, including vision-language-action (VLA) models and world-action models (WAMs). Existing inference runtimes target request-response serving and fail to meet embodied deployment requirements: multi-rate execution within closed-loop control, latency-first batch-1 inference on heterogeneous hardware, and extensible I/O beyond fixed token interfaces. Based on architectural analysis of representative models, Embodied.cpp extracts a shared execution path organized into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. It offers modular multi-rate execution, latency-first fused inference, and extensible operators via a single backend abstraction across devices, robots, and simulators. Benchmarks on HY-VLA and pi0.5 achieved 100.0% and 91.0% task success rates, while a WAM benchmark reduced per-block memory from 312.2 MiB to 88.1 MiB, demonstrating high accuracy with improved deployment efficiency.

Overview

This post summarizes the paper Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Edge Devices (arXiv: 2507.03242), in the computer vision field, by Ling Xu, Chuyu Han, and Borui Li, released 2026-07-04.

Problem

Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment:

  • Multi-rate execution inside closed-loop control
  • Latency-first batch-1 inference on heterogeneous hardware
  • Extensible embodied interfaces beyond fixed token I/O
  • Approach

    Embodied.cpp is a portable C++ inference runtime for embodied models. Based on an architectural analysis of representative VLA models and WAMs, it captures a shared execution path and organizes it into five layers:

    1. Input adapters 2. Sequence builders 3. Backbone execution 4. Head plugins 5. Deployment adapters

    The runtime provides modular multi-rate execution, latency-first fused inference, and extensible operator and I/O support, deployable across heterogeneous devices, robots, and simulators through a single backend abstraction.

    Results

  • Evaluations on two VLA models, HY-VLA and pi0.5, achieved 100.0% and 91.0% task success rates respectively.
  • A preliminary WAM benchmark built on LingBot-VA Transformer blocks reduced block memory from 312.2 MiB to 88.1 MiB.
  • These results indicate that Embodied.cpp significantly improves deployment efficiency while maintaining high accuracy.

    Links

  • Paper: https://arxiv.org/abs/2507.03242

Tags

#embodied-ai#inference-runtime#vla#c-plus-plus#edge-devices#robotics#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208430