A Counterintuitive Finding
Give a multimodal LLM a "magnifying glass"—let it actively crop and zoom into image regions during reasoning. You'd expect it to be more accurate and more careful. The results say otherwise. Across six open-source models and five fine-grained perception benchmarks:
- DeepEyes: 83.3% accuracy on V* with tools on AND off (Δ = 0.0pp); it even *dropped* 1.4pp on HR-Bench-8K.
- Qwen3-VL-4B: Δ = 0.0pp on HR-Bench-4K, Δ = -1.6pp on HR-Bench-8K.
- Mini-o3: +21.3pp on VisualProbe, but with exploding token consumption—and 84.8% of trajectories hit the tool-call limit.
- Paper: https://arxiv.org/abs/2608.06270
- Code: https://github.com/OpenCausaLab/CauAudit
- Observation-mediated path T → O → Y: the observation's visual content genuinely influences the answer ("really looking").
- Action-induced shortcut T → Y: the tool-call text itself drives the answer; O does nothing ("fake looking").
- (a) Saturated prior: the model was already confident before calling (gap g_{i-1} > τ_sat = 0.95). Typical: Qwen3-VL-8B.
- (b) Structurally inactive calls: no call in the trajectory carries any evidence. Typical: DeepEyes.
- (a) Post-saturation continuation: the answer is settled, yet the model keeps calling. Mini-o3's per-position VEG decays while its harmful-call rate rises.
- (b) Budget exhaustion: the model never stops, calling until the cap. Mini-o3's 84.8% Hit-MaxT fits this.
- When evaluating: look at Calibrated share, VEG distributions, Hit-MaxT ratios—not just accuracy.
- When training: give step-level VEG as process reward, not outcome reward alone.
- At inference: bypass tools for Mode 1, early-stop for Mode 2.
In short: the model uses the magnifying glass, spends more tokens, and sometimes answers worse.
This is a structural phenomenon, not an edge case. In the paper *The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images* (arXiv:2608.06270), Zhiheng Wang, Bo Peng, Lai Wei, and Chaochao Lu (Shanghai AI Lab / Shanghai Jiao Tong University / Shanghai Innovation Institute) name it the illusion of visual tool-use.
The Problem: Aggregate Numbers Lie
"Thinking with images" has become a hot paradigm—OpenAI, Qwen, DeepEyes, Pixel Reasoner, Mini-o3, Thyme all insert crop-and-zoom operations into reasoning chains. Benchmarks do go up (Mini-o3: 33.7% → 55.0% on VisualProbe). But the authors ask a sharper question:
> Does the visual evidence obtained from tool calls causally affect the final answer?
Correlation isn't causation—the model may have answered anyway, with tool calls as pure ritual. Answering this requires causal intervention.
Three Levels of Causal Intervention
The authors formalize visual tool-use as a causal graph with two paths:
Level 1: Policy-level (use vs. not use)
Compare accuracy with tools on/off. Gains are limited and highly uneven: DeepEyes ≈ 0; Mini-o3 and Qwen3-VL-8B gain most, but with large token cost. This ATE conflates real and fake looking.
Level 2: Trajectory-level (replace observations with garbage)
Replace every tool-returned observation with a random same-size, same-position crop. Results on V*:
| Model | Acc (orig / intervened) | Δ | Hit-MaxT | |---|---|---|---| | DeepEyes | 83.3 / 83.8 | +0.4 | 0% | | Pixel Reasoner | 84.8 / 82.7 | -2.1 | 0% | | Mini-o3 | 87.8 / 23.6 | -64.2 | 84.8% | | Qwen3-VL-8B | 91.1 / 30.4 | -60.7 | 59.2% | | Qwen3-VL-4B | 86.4 / 38.2 | -48.2 | 50.3% | | Thyme | 83.2 / 82.2 | -1.0 | 0% |
Mini-o3's answers genuinely depend on observations ("really looking"), while DeepEyes' and Thyme's observations contribute almost nothing causally. Note Mini-o3's 84.8% Hit-MaxT: it doesn't "look once and stop"—it looks compulsively.
Level 3: Step-level (contribution of a single observation)
Swap the i-th real observation O_i^real for a counterfactual crop O_i^cf under a fixed prefix and measure the probability change on the correct answer. This is Visual Evidence Gain (VEG):
> VEG_i = ΔM_i^real - ΔM_i^cf
Large VEG means the observation carries real visual evidence; near-zero VEG means the call is "fake looking." VEG is the paper's core tool, separating "calls a tool" from "uses what the tool returns."
Two Failure Modes
Mode 1: Calling Without Looking (CWL)
Tool called, observation irrelevant to the answer. Sub-cases:
> Analogy: a student highlighting an entire textbook without reading the highlighted words. The pen moves; the eyes don't.
Mode 2: Looking Without Planning (LWP)
Observations are informative (nonzero VEG), but the calling plan is broken. Sub-cases:
> Analogy: a doctor who has finished all tests and reached a diagnosis—yet keeps ordering more, until closing time.
Key Finding: A Few "Calibrated" Trajectories Carry All the Gains
Splitting trajectories into CWL, LWP, No-call, and Calibrated:
> Calibrated is the only group with positive ATE on all models, and it accounts for the vast majority of nonzero ATE.
Aggregate gains come almost entirely from a small set of trajectories that genuinely use tools. Most tool-call trajectories are ritual (CWL) or waste (LWP).
> "Having a tool does not necessarily mean it is used properly."
The RL-Trap Hypothesis: Why All Models Fail the Same Way
CWL and LWP appear across six models with different architectures, data, and tool interfaces—suggesting a shared cause. The authors propose the RL-trap hypothesis:
> Outcome-only RL simultaneously produces both failure modes.
1. Reinforces tool-calling behavior: rewarding tool-calls-that-answer-correctly strengthens the T → Y shortcut (Mode 1). 2. Doesn't penalize redundant calls: outcome-only reward can't distinguish harmful vs. redundant intermediate steps (Mode 2). 3. Can't prefer calibrated trajectories: a Calibrated trajectory and a Mode 1 trajectory that both answer correctly get identical rewards.
If true, the flaw is in the training objective itself—not model size or tool quality.
Engineering Value: The Diagnosis Is Actionable
1. Evaluate beyond accuracy: track distribution shifts—does the Calibrated fraction Δf_Cal rise? Do CWL + LWP fractions fall? ATE can stay flat while genuine improvement happens (and vice versa). 2. Inference-time adaptation: bypass tool calls for Mode 1 trajectories (saves tokens); apply early-stopping for Mode 2. 3. Process-aware credit assignment: use step-level VEG as the per-step supervision that outcome-only reward lacks—reward Calibrated steps, penalize wasted and harmful calls.
Commentary: The Evaluation Blind-Spot Law
The post's author connects this to a recurring theme—the evaluation blind-spot law: models optimize what you measure; problems hide where you don't. Examples cited include RLHF-induced self-correction artifacts, token-budget bimodality invisible in standard loss, quantization bias under standard safety checks, and agent skill libraries that look good on average pass rates while masking regressions. This paper adds a precise strike: aggregate accuracy is the most-used metric and the biggest blind spot, mixing "a few trajectories truly use tools" with "most calls are ritual or waste." The remedy points the same direction: from aggregate to distributional metrics, from outcome to process metrics.
A related principle—granularity isomorphism: the optimization granularity should match the object being optimized. The working unit of thinking-with-images is the *individual tool call*, but outcome-only RL rewards the *whole trajectory*. Any RL setting where reward granularity is coarser than behavior granularity leaves room for "fake doing" and "busy flailing": code agents, search agents, long reasoning chains alike.
Limitations and Open Questions
1. Model access: experiments are on open-source models only; step-level VEG needs token-level probabilities, so closed models (OpenAI o3/o4-mini) couldn't be audited. 2. Limited toolset: only crop-and-zoom tested; OCR, segmentation, video frame selection, and external search may differ. 3. RL-trap is a hypothesis: establishing the mechanism requires matched training experiments varying only the reward signal—left to future work.
Conclusion: From "Can Call Tools" to "Actually Uses Tools"
This paper turns an intuitive suspicion into an operational causal framework (three intervention levels + VEG) and exposes a structural problem across six models. The most striking finding: aggregate accuracy rises, but the gains concentrate in a few calibrated trajectories—most tool calls are ceremony or waste. The actionable takeaways:
---
Paper: The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images Authors: Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu Institutions: Shanghai AI Lab / Shanghai Jiao Tong University / Shanghai Innovation Institute arXiv: https://arxiv.org/abs/2608.06270 Code: https://github.com/OpenCausaLab/CauAudit License: CC BY 4.0