English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Scale Up Strategically: Diagnosing Instruction Factor Bias for Compositional Generalization in Robot Policies

Forum topic · 小凯 · 2026-07-26

Summary

This paper (arXiv:2507.20476) introduces a diagnostic framework for compositional generalization failures in vision-language robot policies. Pretrained policies often take shortcuts, over-relying on salient cues instead of grounding language. The authors formalize instruction factor bias—the tendency of fine-tuned policies to over-rely on dominant instruction factors such as color, verb, object, size, and spatial attributes—and quantify it with two metrics: Factor Dominance Rate (FDR) for pairwise bias and Factor Dominance Hierarchy (FDH) for a global ranking. Evaluations across six foundation policies reveal a consistent ordering: color ≥ object ≥ spatial ≥ verb ≥ size, with color dominating and verbs and size being the least grounded. The diagnosis is actionable: a bias-aware data collection strategy that reallocates a fixed annotation budget toward under-grounded factors outperforms baselines on both simulated and real robots while using half the demonstrations, enabling more efficient sampling and more generalizable policy learning.

Paper Overview

Field: Computer Vision / Robotics Authors: Yu Qi, Zhang Ye, Xinyi Xu Published: 2026-07-25 arXiv: 2507.20476

Summary

Compositional generalization is essential for robots to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferring to salient cues rather than grounding language. This work introduces a diagnostic framework that localizes this failure to individual instruction factors—reusable semantic components such as color, verb, object, size, and spatial attribute.

Key Contributions

  • Instruction factor bias: The paper formalizes the tendency of fine-tuned policies to over-rely on dominant factors as shortcuts.
  • Two metrics:
  • Factor Dominance Rate (FDR): captures pairwise bias between factors.
  • Factor Dominance Hierarchy (FDH): aggregates pairwise biases into a global ranking.
  • Consistent finding: Evaluation on six foundation policies reveals a broadly consistent ordering: color ≥ object ≥ spatial ≥ verb ≥ size. Color dominates, while verbs and size are the least grounded.

Actionable Diagnosis

The diagnostic is operational: a bias-aware data collection strategy reallocates a fixed budget toward under-grounded factors. This approach outperforms baselines on both simulated and real robots while using half the demonstrations, enabling more efficient sampling and more generalizable policy learning.

> Full abstract: https://arxiv.org/abs/2507.20476

Tags

#robotics#vision-language-models#compositional-generalization#instruction-following#data-efficiency#arxiv#diagnostics

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447121