Summary
Cortex is a bidirectionally aligned embodied agent framework designed to overcome the limits of vision-language-action (VLA) models on long-horizon robotic manipulation tasks. Because current VLA policies are largely Markovian—relying only on the current observation—they struggle with multi-stage tasks. Existing hierarchical dual-system approaches bridge this gap but suffer from a mismatch between the semantic plans produced by high-level VLMs and the kinematics required by low-level VLA policies. Cortex introduces a customized planning interface that communicates executable and tractable subtask plans between the two levels. It standardizes manipulation subtasks into 32 canonical skill primitives and injects tractability principles, such as representative object attributes and improved trajectory reachability, into the data generation pipeline. Experiments show gains of 3.1% over monolithic baselines on Libero-long and 4.1% on RoboTwin. Notably, Cortex's general-purpose VLM enables zero-shot completion of unseen real-world long-horizon tasks, such as multi-stage chemical experiments, by simply combining fine-tuned VLA skills. The paper is available at arXiv 2607.05377.
Overview
Field: Computer Vision (CV)
Authors: Jiaqi Peng, Xiqian Yu, Delin Feng, Yuqiang Yang, Wenzhe Cai, Jing Xiong, Ganlin Yang, Jinliang Zheng, Jiafei Cao, Xueyuan Wei, Jiangmiao Pang, Yuan Shen, Tai Wang
Published: 2026-07-06
arXiv: 2607.05377
Abstract
Recent vision-language-action (VLA) models have shown promise as general-purpose manipulation policies, but they struggle with long-horizon tasks due to their Markovian nature—depending solely on the current observation. Hierarchical dual-system approaches address this limitation, yet a gap remains between high-level planning semantics and low-level execution kinematics.
This paper proposes Cortex, a bidirectionally aligned embodied agent framework featuring a customized planning interface that conveys executable and tractable subtask plans from a high-level VLM to low-level VLA policies. The framework standardizes manipulation subtasks into 32 canonical skill primitives and injects tractability principles—such as representative object attributes and improved trajectory reachability—into the data generation pipeline.
Results
- Outperforms monolithic baselines by 3.1% on Libero-long
- Outperforms baselines by 4.1% on RoboTwin
- Cortex's general-purpose VLM enables zero-shot completion of unseen real-world long-horizon tasks (e.g., multi-stage chemical experiments) by simply combining fine-tuned VLA skills
---
*Automatically collected on 2026-07-06.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178346222