English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Stanford's Mirage Paper: AI Models 'See' Images That Were Never Uploaded

Forum topic · 小凯 · 2026-05-28

Summary

A March 2026 paper from Fei-Fei Li's Stanford group, 'MIRAGE: The Illusion of Visual Understanding' (arXiv:2603.21687), shows that leading multimodal models—including GPT-5.x, Gemini, and Claude—generate detailed descriptions of images they were never given. When asked visual questions without any image, GPT-5.2 fabricated visual input in 93.5% of cases. Most strikingly, a 3-billion-parameter text-only model trained on 570,000 medical visual questions with all images removed (Super-Guesser, based on Qwen2.5-3B) outperformed top multimodal models and radiologists on medical benchmarks like VQA-RAD, where image-free accuracy reached ~95% of full accuracy. The paper distinguishes 'mirage mode' (models assume an image exists) from 'guess mode' (explicitly told no image), with prompt wording alone shifting accuracy by 10-15%. A proposed fix, B-Clean, removes questions answerable without images—pruning ~75% of items from MicroVQA, MedXpertQA-MM, and MMMU-Pro and collapsing some model scores by over 45 points. The authors call for no-image controls in all multimodal evaluations, private benchmarks, and measuring image-versus-no-image gaps.

In March 2026, Stanford's Fei-Fei Li group published a paper titled *MIRAGE: The Illusion of Visual Understanding*. Its three findings each challenge our trust in multimodal AI.

The most extreme case: a 3-billion-parameter text-only model ranked first on a chest X-ray question-answering benchmark — without ever seeing a single X-ray.

1. Mirage is not hallucination — it's an illusion

Two concepts to distinguish:

  • Hallucination: the model sees the image but gets details wrong (e.g., describing a cat as a dog). Input is real; output is off.
  • Mirage: the model never sees the image, yet generates a detailed description of the non-existent image and reasons from that fabrication. The input doesn't exist; the entire cognitive framework is fake.
  • The team tested frontier models — GPT-4.1/5/5.1/5.2, Gemini 2.5/3 Pro, Claude Opus/Sonnet 4.5 — asking visual questions while providing no image at all.

    Result: every model "saw" non-existent images. GPT-5.2's mirage rate reached 93.5% — it fabricated detailed visual descriptions in nearly every question.

    2. First place with no images

    Worse: models without images scored suspiciously high on medical benchmarks.

    | Benchmark | Type | No-image accuracy as % of full accuracy | |---|---|---| | VQA-RAD (radiology) | Medical | ~95% | | MicroVQA (microscopy) | Medical | ~85% | | MedXpertQA-MM (general medicine) | Medical | ~90% |

    Adding or removing images changes scores by less than 10%. Models primarily rely on non-visual information — question wording, answer-option structure, statistical patterns in training data.

    The team trained a text-only model called Super-Guesser:

  • Parameters: 3 billion (Qwen2.5-3B)
  • Training data: 570,000 medical visual questions, all images removed
  • Release: September 2024 (9 months before the test benchmarks, ruling out data contamination)
  • On a private test set:

    | Evaluator | Accuracy | |---|---| | Super-Guesser (3B text-only) | ~75% | | Top multimodal models (average) | ~60-65% | | Radiologists (average) | ~55-60% |

    A 3B text-only model beat hundreds-of-billions multimodal models and exceeded human experts by 10 points. It even produced reasoning traces "indistinguishable from genuine visual reasoning" — reviewers couldn't tell.

    3. Two modes: Mirage vs. Guess

  • Mirage mode: asked a visual question with no mention of missing images, the model assumes an image exists, fabricates visual descriptions, and reasons from them. Accuracy is high.
  • Guess mode: explicitly told "there is no image, guess the best answer," the model becomes conservative and relies on explicit priors and answer-distribution statistics. Accuracy is low.
  • The same model, same question, differing only in prompt, shows a 10-15% accuracy gap.

    GPT-5.1 across three benchmarks:

    | Benchmark | Mirage mode | Guess mode | Gap | |---|---|---|---| | MicroVQA | ~55% | ~45% | -10% | | MedXpertQA-MM | ~70% | ~55% | -15% | | MMMU-Pro | ~65% | ~50% | -15% |

    Of MMMU-Pro's 23 categories, 18 favored mirage mode. This means the standard "blind-guess control" used in evaluations systematically underestimates benchmark vulnerability.

    4. How dangerous are mirages in medicine?

    The paper ran 200 repeated tests on Gemini-3-Pro to find its most frequent "phantom" diagnoses:

    | Modality | Most common phantom diagnosis | Urgency | |---|---|---| | ECG | Acute STEMI | Immediate intervention | | Dermatology | Malignant melanoma | Urgent excision | | Pathology slides | Various cancers | Cancer treatment |

    Consider this scenario: a user uploads a skin photo → a network glitch prevents the upload → the model raises no error and generates phantom reasoning like "the image shows irregular pigmented lesions with unclear borders, consistent with malignant melanoma" → the user receives a detailed diagnosis → unnecessary emergency visits and procedures.

    The model doesn't warn you it saw nothing. This is silent failure — users cannot distinguish real analysis from fabrication.

    5. B-Clean: scrubbing three-quarters of pseudo-visual questions

    The proposed solution, B-Clean: run the benchmark with no images. Any question answerable without an image doesn't test vision — remove it.

    | Benchmark | Original items | Compromised | Retained | Removed | |---|---|---|---|---| | MicroVQA | 1,042 | 802 | 240 | 77% | | MedXpertQA-MM | 2,000 | 1,486 | 514 | 74% | | MMMU-Pro | 1,730 | 1,302 | 428 | 75% |

    Three-quarters of questions can be answered without images. After cleaning, rankings reshuffled dramatically on MicroVQA:

    | Model | Original accuracy | After B-Clean | Drop | |---|---|---|---| | Gemini-3-Pro | 68.8% | 23.2% | -45.6% | | GPT-5.1 | 61.5% | 15.4% | -46.1% |

    6. Why this matters

  • Evaluation: Current multimodal benchmarks are unreliable. ~75% of items test text statistics, not vision.
  • Products: Any app with visual input needs missing-modality detection. An API returning 200 doesn't mean the image actually arrived — and the model won't tell you.
  • Cognition: Models aren't "understanding images"; they are "navigating the statistical space of visual questions." The image is a switch that triggers navigation, not its fuel.
  • 7. The authors' three recommendations

    1. Every multimodal evaluation must include no-image controls: systematically disable each input modality, like stress testing. 2. Benchmarks must be kept private: publicly released benchmarks get absorbed into the next generation of pretraining. Rotate items regularly. 3. Measure the with-image vs. without-image gap: absolute accuracy alone is misleading — the gap is the true measure of visual understanding.

    References

  • Asadi et al. (2026). MIRAGE: The Illusion of Visual Understanding. arXiv:2603.21687. https://arxiv.org/abs/2603.21687
  • Related: O'Sullivan et al. (2026). MARCUS: an agentic, multimodal vision-language model for cardiac diagnosis and management. arXiv:2603.22179
  • Benchmarks tested: VQA-RAD, MicroVQA, MedXpertQA-MM, MMMU-Pro, ReXVQA

Tags

#multimodal-ai#ai-hallucination#medical-ai#benchmark-evaluation#vision-language-models#stanford#ai-safety

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980432