Overview
Field: ML Authors: Maggie Wang, Lars Osterberg, Stephen Tian Published: 2026-06-24 arXiv: 2506.14748
Key Points
Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. InSight addresses this by unlocking autonomous skill acquisition through primitive-action-level steerability.
- Primitive steerability: VLAs become steerable at the level of primitive actions (e.g., "move gripper to the bowl", "lift upward", "pour the bottle").
- Automated segmentation pipeline: Demonstrations are partitioned into labeled primitives via VLM plan decomposition combined with end-effector poses, enabling VLA primitive steerability.
- VLM-guided data flywheel: The system identifies missing primitives required for a novel task, autonomously attempts demonstrations using VLM-proposed low-level control, and automatically labels, stores, and integrates successful demonstrations into the VLA training set.
- Evaluation: Tested on simulated and real-world manipulation tasks — flipping a block, closing a drawer, sweeping, twisting, and pouring — with no human demonstrations of the target skills.
- Composition: Once learned, primitives can be composed to perform new long-horizon tasks without additional human demonstrations.
Conclusion
The findings suggest that primitive-action steerability provides a practical foundation for continual skill acquisition in VLA policies.
--- *Auto-collected on 2026-06-25*