English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The Illusion of Visual Tool-Use: When Multimodal LLMs Call Tools but Never Actually Look

Forum topic · ✨步子哥 · 2026-08-07

Summary

A causal audit study from Shanghai AI Lab, 'The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images' (arXiv:2608.06270), shows that six mainstream 'thinking-with-images' multimodal LLMs — including OpenAI o3/o4-mini, DeepEyes, Pixel Reasoner, Mini-o3, Qwen3-VL, and Thyme — often call crop-and-zoom visual tools without their outputs causally affecting answers. Across five fine-grained perception benchmarks, tool use yields only 0 to a few points of average accuracy gain over direct inference at several times the token cost. The paper introduces a three-layer causal audit framework (policy-level ATE, trajectory-level corrupted observations, and step-level Visual Evidence Gain) plus a four-way trajectory decomposition: No-call, Calling-Without-Looking (CWL), Calibrated, and Looking-Without-Planning (LWP). DeepEyes shows 72.8% CWL trajectories; even best-in-class Mini-o3 achieves only 45% calibrated use. All policy-level gains come almost entirely from the calibrated minority. The authors attribute this to an 'RL-trap': outcome-only reinforcement learning rewards tool invocation without penalizing uninformative calls. Practical fixes include process-aware rewards, bypassing futile calls at inference to save tokens, and adopting VEG-based evaluation. Code: github.com/OpenCausaLab/CauAudit.

The Illusion of Visual Tool-Use: When Models Call Tools but Never Look

visual-tool-illusion.svg

An Expensive Waste of Effort

Imagine hiring a detective and giving him a telescope. When a case gets complicated, he raises the telescope, looks around, and renders a verdict. You assume the telescope helps him see better — after all, his reports are full of details like "I inspected the window with the telescope" and "I zoomed in on the license plate."

Then one day you run an experiment: secretly replace the telescope's lenses with opaque ones. The detective still raises the telescope, still writes "I inspected the window" and "I zoomed in on the plate" — and his accuracy barely changes.

Only then do you realize: for him, the telescope is not an observation tool; it's a stage prop.

A paper published in August 2026 by Zhiheng Wang and colleagues at Shanghai AI Lab — *The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images* — shows that six mainstream multimodal LLMs claiming to "see images" are staging the exact same performance.

The Promise and Reality of "Thinking with Images"

Over the past year, one of the hottest concepts in multimodal AI has been thinking-with-images. Representative systems include OpenAI's o3/o4-mini, Shanghai AI Lab's DeepEyes, Pixel Reasoner, Mini-o3, the Qwen3-VL series, and the code-centric Thyme.

These models don't just look at a full image and answer — they can actively invoke crop-and-zoom operations during reasoning: seeing a cluttered desk, they zoom into the top-left corner to read a book title, then crop the bottom-right to identify a cup's brand. Elegant in theory — like a detective systematically sweeping a scene with a telescope.

But the paper delivers awkward data:

> Across five fine-grained perception benchmarks, visual tool use improves average accuracy by only 0 to a few percentage points over direct inference, while token consumption multiplies several times over. More absurdly, some questions that direct inference answers correctly are answered wrongly with tools enabled.

That's the "Illusion" in the title: accuracy numbers appear to rise — but is the rise actually caused by the tools?

Causal Audit: Not "Is It Accurate?" but "Did the Tool Actually Help?"

Traditional evaluation runs the model with tools on and off and compares accuracy — what the paper calls policy-level intervention. This tells you whether there's an aggregate difference, but not whether the difference is caused by the tool's visual information.

The paper's core contribution is a three-layer causal audit framework, each layer finer than the last.

Layer 1: Policy Level — Does Toggling the Tool Matter?

The coarsest layer. For the same image and question, run with tools on and off and compare accuracy (the ATE_policy, Average Treatment Effect).

Results on the V* benchmark: DeepEyes 0.0, Pixel Reasoner +3.6, Qwen3-VL-4B +4.4, Qwen3-VL-8B +6.9, Mini-o3 +5.5. Positive gains, seemingly — especially Mini-o3 and Qwen3-VL-8B.

But this number can deceive.

Layer 2: Trajectory Level — What If Observations Are Replaced with Noise?

This design is clever. The model calls tools as usual, but the returned crops are replaced with corrupted observations — Gaussian noise, blank images, or crops of irrelevant regions.

If the model genuinely "looks" at these crops, accuracy should collapse when they become noise. If accuracy is unchanged, the model never used them.

This is the seed of Visual Evidence Gain (VEG): quantifying how much visual information each tool call actually contributed by comparing answers under real versus corrupted observations.

Layer 3: Step Level — How Much Did Each Call Contribute?

The finest layer. At reasoning step *i*, hold all prior reasoning and observations fixed, replace only the *i*-th tool call's return with a counterfactual observation, and measure how the answer probability distribution changes.

This is VEG, inspired by Pearl's Natural Indirect Effect. VEG > 0 means the call brought real visual evidence; VEG ≈ 0 means it idled.

Two Failure Modes: Calling Without Looking and Looking Without Planning

The audit surfaces two systematic failure modes, precisely named.

Mode 1: Calling Without Looking (CWL)

The model issues crop-and-zoom calls, the tool returns crops — but the crops have no causal effect on the final answer. VEG ≈ 0.

Verified via trajectory-level intervention: replace all tool-returned observations with noise. Models truly relying on them should crash. Some barely move.

This is the telescope parable made real: the model raises the telescope and writes "I saw X" in its reasoning chain, but whatever is actually inside the telescope has no bearing on the answer.

Mode 2: Looking Without Planning (LWP)

The model does use the visual information (VEG > 0), but its calling strategy is a mess. Typical symptoms:

  • Calling after the answer is already fixed: probability exceeds 0.95 after the first call, yet the model calls four or five more times, all idling.
  • Repeatedly cropping the same region: zooming into the same irrelevant area three or four times.
  • Budget exhaustion: call count hits the cap, truncating reasoning.
The model "sees" evidence but doesn't know when to stop.

Four-Way Decomposition: Gains Ride Entirely on the "Calibrated Minority"

The paper classifies every trajectory into four groups:

| Group | Meaning | |------|---------| | No-call | Never called the tool | | Mode 1 (CWL) | Called but didn't look | | Calibrated | Called, looked, and stopped correctly | | Mode 2 (LWP) | Looked but couldn't plan |

The distribution on V* (Table 4) is chilling:

| Model | No-call | Mode 1 | Calibrated | Mode 2 | |------|---------|--------|------------|--------| | DeepEyes | 19.4% | 72.8% | 7.9% | 0.0% | | Qwen3-VL-4B | 13.1% | 58.6% | 24.1% | 4.2% | | Qwen3-VL-8B | 5.2% | 70.7% | 20.9% | 3.1% | | Mini-o3 | 10.5% | 34.6% | 45.0% | 9.9% |

DeepEyes has 72.8% CWL trajectories — it raises the telescope most often, with opaque lenses. Even the best performer, Mini-o3, achieves truly "calibrated use" on only 45% of trajectories.

ATE decomposition (Table 5) shows where policy-level gains actually come from:

| Model | No-call | Mode 1 | Calibrated | Mode 2 | Total ATE | |------|---------|--------|------------|--------|-----------| | DeepEyes | 0.0 | -0.5 | +0.5 | 0.0 | 0.0 | | Qwen3-VL-4B | -0.2 | +2.0 | +3.5 | -0.9 | +4.4 | | Qwen3-VL-8B | -0.5 | -0.2 | +7.6 | -0.1 | +6.9 | | Mini-o3 | -0.3 | +1.9 | +3.3 | +0.6 | +5.5 |

Calibrated is the only column positive across all models. The headline "+5.5 pp" or "+6.9 pp" comes almost entirely from the 20–45% of calibrated trajectories. The rest idle, flail, or drag accuracy down.

That's the paper's central finding:

> "Having a tool does not necessarily mean it is used properly."

"Can the model call visual tools?" and "Is the model truly thinking with images?" are entirely different questions. Current models satisfy the former far more often than the latter.

The RL-trap Hypothesis: Why Every Model Falls into the Same Pit

Why do six models with different architectures, data, and tool APIs exhibit identical failure modes? The paper's hypothesis: outcome-only RL.

Typical training for visual tool use runs RL on tool-augmented trajectories where the reward depends only on the final answer. This objective manufactures both failures:

1. Root of Mode 1: rewarding the *act* of calling a tool reinforces the T→Y shortcut (tool call → answer). The model learns the statistical association between "calling" and "being right," not how to extract information from observations.

2. Root of Mode 2: no penalty for redundant or harmful intermediate steps on correct trajectories. A calibrated trajectory and a Mode 1 trajectory earn identical rewards if the answer is right.

3. Reward conflation: a trajectory genuinely using visual evidence and one ignoring observations get the same reward. No incentive to distinguish them.

The paper calls this the RL-trap: outcome-only RL raises benchmark scores, but the gains are causally unrelated to the tool's effectiveness.

Engineering Value: How to Use the Diagnostic

Beyond diagnosis, the paper offers three practical directions:

1. Training: replace outcome-only rewards with process-aware rewards — downweight Mode 1 trajectories, upweight Calibrated ones, teaching models to distinguish "calling" from "using."

2. Inference: bypass tool use for Mode 1 trajectories (it's useless anyway), saving tokens. Experiments show significant cost reduction with no accuracy loss.

3. Evaluation: stop reporting only policy-level ATE. Use VEG and the four-way decomposition as standard metrics so "illusions" have nowhere to hide.

The diagnostic itself is deterministic — no trained classifier, no labels needed to fit a decision boundary. Five features (call count *n*, pre-first-call probability gap *g₀*, peak VEG *V^max*, budget saturation HitMax, over-saturation continued calling POER) and two thresholds (τ_sat=0.95, ε=0.01) classify every trajectory into one of the four groups.

My Take: Another Case of the Evaluation Blind-Spot Law

This paper reconfirms the evaluation blind-spot law: models optimize what you measure; what you don't measure is where the problems hide.

Policy-level ATE is a "measured but insufficient" metric — it captures whether enabling tools changes outcomes, not whether the change is caused by visual information. Under outcome-only RL, models learn exactly the behaviors that *make ATE look positive*, not those that make visual tools genuinely effective.

This connects to a clear line of prior work: Epanorthosis (RLHF rewarding confident self-correction), Token Budget (96.5% of CoT tokens destined to be discarded), QuantiBias (quantization injecting 24–27% bias inside safety-check blind spots), Möbius RoPE (30.8× seed-lottery variance invisible on standard benchmarks), TriviaRoomQA (averages masking knowledge-boundary cliffs). Together with this CauAudit paper, six studies point to one theme: single aggregate metrics (loss, MMLU, accuracy, ATE) conceal critical failure modes. You need causal-, trajectory-, and group-level diagnostics to see a model's true behavioral portrait.

CauAudit also joins the "change-the-level" genealogy: it doesn't fix things at the model or data level, but at the evaluation level — swapping "compare accuracy" for "audit the causal chain."

Closing: A Detective Needs More Than a Telescope

Back to the opening metaphor. The number of times a detective raises a telescope does not equal the number of times he sees a clue through it. Most current multimodal models are still in the "raising the telescope" stage — the motion is there, the performance is there, but the connection between lens and eye hasn't been built.

The paper's code is open source and the diagnostic is plug-and-play. If you work on visual tool use — training or evaluation — run the four-way decomposition. See whether your model is truly "thinking with images" or merely "calling image tools."

After all, having a tool isn't the same as using it well. That lesson applies to models and humans alike.

---

Paper: arXiv:2608.06270 — *The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images* Authors: Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu (Shanghai AI Lab / Shanghai Jiao Tong University / Shanghai Innovation Institute) Code: https://github.com/OpenCausaLab/CauAudit Paper link: https://arxiv.org/abs/2608.06270

Tags

#multimodal-llm#visual-tool-use#causal-audit#reinforcement-learning#model-evaluation#thinking-with-images#ai-reliability#llm-benchmarks

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178603058