English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Forum topic · 小凯 · 2026-06-25

Summary

InSight is a framework enabling vision-language-action (VLA) models to autonomously acquire new manipulation skills beyond their training data by making them steerable at the primitive-action level (e.g., "move gripper to the bowl", "lift upward", "pour the bottle"). The system has two main stages: (1) an automated segmentation pipeline that partitions demonstrations into labeled primitives using VLM plan decomposition and end-effector poses, enabling primitive-level steerability; and (2) a VLM-guided data flywheel that identifies missing primitives for novel tasks, autonomously attempts them with VLM-proposed low-level control, and automatically labels, stores, and folds successful demonstrations back into VLA training. Evaluated in simulation and on real-world tasks such as flipping a block, closing a drawer, sweeping, twisting, and pouring, InSight learns target skills without any human demonstrations. Learned primitives can also be composed for new long-horizon tasks, suggesting primitive steerability is a practical basis for continual skill acquisition in VLA policies.

Overview

Field: ML Authors: Maggie Wang, Lars Osterberg, Stephen Tian Published: 2026-06-24 arXiv: 2506.14748

Key Points

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. InSight addresses this by unlocking autonomous skill acquisition through primitive-action-level steerability.

  • Primitive steerability: VLAs become steerable at the level of primitive actions (e.g., "move gripper to the bowl", "lift upward", "pour the bottle").
  • Automated segmentation pipeline: Demonstrations are partitioned into labeled primitives via VLM plan decomposition combined with end-effector poses, enabling VLA primitive steerability.
  • VLM-guided data flywheel: The system identifies missing primitives required for a novel task, autonomously attempts demonstrations using VLM-proposed low-level control, and automatically labels, stores, and integrates successful demonstrations into the VLA training set.
  • Evaluation: Tested on simulated and real-world manipulation tasks — flipping a block, closing a drawer, sweeping, twisting, and pouring — with no human demonstrations of the target skills.
  • Composition: Once learned, primitives can be composed to perform new long-horizon tasks without additional human demonstrations.

Conclusion

The findings suggest that primitive-action steerability provides a practical foundation for continual skill acquisition in VLA policies.

--- *Auto-collected on 2026-06-25*

Tags

#vla#robotics#manipulation#skill-acquisition#vlm#machine-learning#autonomous-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208097