Actions Speak Louder Than Words: 2.38M Agent Rollouts Reveal How LLM Agents Differ Across Languages
> Paper: arXiv:2608.11110 > Code: github.com/souro/agent-actions-speak-louder-than-words > Authors: Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram (Microsoft Research India)
A Question That Keeps Agent Engineers Up at Night
Ask a tool-calling agent: "Check tomorrow's weather, then book a 7 PM restaurant." In English, it calls three APIs perfectly. What does it do in Hindi?
Before this paper, nobody had really answered that.
Existing multilingual benchmarks compare final answers (e.g., MGSM, xCopA) or chain-of-thought text quality. But an agent's core behavior is not what it *says*—it's what it *does*: which tools it calls, in what order, with which parameters.
A correct final answer doesn't mean correct intermediate steps. The agent might translate Hindi to English, query the weather, and translate back—introducing extra steps, failure points, and latency.
The Scale: 2.38 Million Rollouts
- 8 models: from GPT-OSS-120B down to Qwen2.5-7B
- 6 benchmarks: tool calling, API operations, multi-step reasoning
- 41 languages: from high-resource (Chinese, Spanish) to extremely low-resource (Kinyarwanda, Telugu)
- 2.38M agent rollouts
- Action policy as a first-class metric: Task success rate hides process differences. Two agents can both succeed—one in 3 steps, one in 7. In production, extra steps mean extra failure points, latency, and cost.
- Causal evidence for the English pivot**: Prior work found latent English representations (Wendler et al. 2024; Zhao et al. 2024) as correlation. This paper tests it causally at the action level. The bottleneck for multilingual agents isn't comprehension—it's whether they can reason *without* English. Currently, they can't.
- Evaluation blind spots: Every uncorrected confounder is a blind spot, and the bias direction is systematically *underestimating* the problem.
Five Confounders That Fool Naive Measurement
1. Short traces look more consistent — trivially, because there's little to differ. 2. Empty traces are perfectly consistent — total failure scores 100% similarity. 3. Random traces exceed 50% consistency — overlapping common APIs inflate the baseline. 4. Self-reproducibility is a ceiling — same model, task, language, prompt, run twice: traces still differ. This is the model's noise floor. 5. Same language, same question, still inconsistent — so how much of a cross-language gap is really about language?
Counterintuitive Result: Every Correction Makes the Effect Bigger
Normally, removing confounders shrinks an effect. Here, **each of the five corrections *increased* the measured cross-lingual difference—because the confounders were all creating false-positive consistency that masked real systematic gaps.
Core Findings
1. Structural, not sampling noise: Under greedy decoding (temperature=0), cross-lingual differences were positive in all 24 test cells. 2. 71-73% policy retention: For frontier models (GPT-OSS-120B, Claude, Gemini), English action policies survive translation only ~71-73%—roughly 1 in 4 actions differs. 3. Model identity explains only 5.7% of variance: Which language you use matters far more than which model you pick. 4. Below 10B parameters, regularity collapses: Small models aren't just worse—they're unpredictable on multilingual agent tasks. 5. The English pivot is causally load-bearing: Agents process non-English tasks by internally translating to English, reasoning in English, and translating back. Instructing models to avoid English reasoning succeeded less than 1% of the time. 6. A regex created a false positive: GPT-OSS-120B's apparent multilingual failure was caused by a regex bug in the evaluation code. You may think you're measuring the model, but you're actually measuring your regex.
Why It Matters
Practical Takeaways
For multilingual agent builders: 1. Don't test only final answers—add a corrected action-policy consistency metric. 2. Treat the English pivot as reality: make translation steps explicit, controlled, and monitorable. 3. Don't build multilingual agents below 10B parameters.
For evaluation designers: 1. Audit your regexes—every one is a potential false-positive source. 2. Report corrections for all five confounders. 3. Report greedy-decoding results to eliminate sampling noise.
Honest Limitations
1. Only tool-calling agents were tested; conversational or planning agents may differ. 2. Language coverage is biased—African and Indigenous languages are underrepresented. 3. The paper proves the English pivot's causality but offers no way to break it—an open problem. 4. 2.38M rollouts is industrial-scale; hard for academic teams to reproduce.
One-Sentence Summary
When you talk to an agent in another language, it appears to think in that language—but it's secretly translating to English, reasoning in English, and translating back. 2.38M experiments show this English pivot is not a habit but a causal necessity: forbid it, and the model simply fails.
> Paper: Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents > Code: github.com/souro/agent-actions-speak-louder-than-words > Institution: Microsoft Research India | Published: 2026-08-11