论文概要
研究领域: CV
作者: Boyao Han, Chen Shi, Jingjing Qian, ZhuoTan Tian, Li Jiang
发布时间: 2026-10-08
arXiv: 2610.12451
中文摘要
视觉-语言-动作(VLA)模型已成为机器人操作的强大基础,但训练时对固定相机配置的依赖使其在部署中面对相机数量或位姿变化时十分脆弱。为克服这些局限,我们提出 VersaCamVLA——一个相机可配置的框架,将相机组表征与动作学习解耦。VersaCamVLA 学习统一的场景 token 接口,将任意可变数量的带位姿 RGB 视图映射为固定尺寸的隐式场景 token,通过多信号目标视图预测和腕部增强位姿采样(WAPS)实现,后者利用自然的腕部相机运动获得免费的位姿多样性。部署时,轻量级空间编码器将这些紧凑场景 token 注入预训练基础 VLA 作为补充视觉条件,无需显式 3D 感知或新视角渲染。在 RoboTwin、LIBERO 和真实机器人平台上的实验表明,VersaCamVLA 持续优于先前 VLA 方法和直接多视图基线,在不同相机数量和未见过的相机位姿下均保持鲁棒的性能。
原文摘要
Vision-Language-Action (VLA) models have emerged as powerful foundations for robotic manipulation, but their reliance on fixed camera configurations during training makes them brittle to changes in camera count or pose during deployment. To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. VersaCamVLA learns a unified scene-token interface that maps an arbitrary, variable set of posed RGB views into fixed-size latent scene tokens. This is achieved via multi-signal target-view prediction and Wrist-Augmented Pose Sampling (WAPS), which leverages natural wrist-camera motion for free pose diversity. At deployment, a lightweight spatial encoder injects these compact scene tokens into a pretrained ba...
自动采集于 2026-10-10
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。