English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CapVector: Learning Transferable Capability Vectors for Efficient VLA Model Finetuning

Forum topic · 小凯 · 2026-05-13

Summary

CapVector (arXiv:2505.07230) is a novel finetuning approach for pretrained Vision-Language-Action (VLA) models, authored by Wenxuan Song, Han Zhao, and Fuhao Li, published May 9, 2025. The paper addresses a key limitation of standard supervised finetuning (SFT), which often fails to effectively improve performance or reduce adaptation costs. While advanced finetuning methods using auxiliary training objectives can boost performance and cut convergence steps, they typically impose significant computational overhead from the extra auxiliary losses. CapVector learns transferable capability vectors to simultaneously enhance finetuning performance and reduce adaptation cost, avoiding the heavy computation of auxiliary objectives. This makes it a practical alternative for adapting large VLA models in robotics and computer vision applications.

Paper Overview

Research area: Computer Vision (CV) Authors: Wenxuan Song, Han Zhao, Fuhao Li Published: 2025-05-09 arXiv: 2505.07230

Abstract

This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adaptation costs during standard supervised finetuning (SFT). Some advanced finetuning methods with auxiliary training objectives can improve performance and reduce the number of convergence steps. However, they typically incur significant computational overhead due to the additional losses from auxiliary objectives. To simultaneously achieve the enhancement of performance and reduction of adaptation cost, the authors introduce CapVector, a method that learns transferable capability vectors in parametric space.

*(Note: the source abstract is truncated; see the arXiv page for the full text.)*

Links

  • arXiv: https://arxiv.org/abs/2505.07230
---

*Auto-collected post from zhichai.net, originally published 2026-05-13.*

Tags

#vla-models#finetuning#supervised-finetuning#computer-vision#machine-learning#arxiv#capability-vectors#robotics

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619927