In March 2026, Stanford's Fei-Fei Li group published a paper titled *MIRAGE: The Illusion of Visual Understanding*. Its three findings each challenge our trust in multimodal AI.
The most extreme case: a 3-billion-parameter text-only model ranked first on a chest X-ray question-answering benchmark — without ever seeing a single X-ray.
1. Mirage is not hallucination — it's an illusion
Two concepts to distinguish:
- Hallucination: the model sees the image but gets details wrong (e.g., describing a cat as a dog). Input is real; output is off.
- Mirage: the model never sees the image, yet generates a detailed description of the non-existent image and reasons from that fabrication. The input doesn't exist; the entire cognitive framework is fake.
- Parameters: 3 billion (Qwen2.5-3B)
- Training data: 570,000 medical visual questions, all images removed
- Release: September 2024 (9 months before the test benchmarks, ruling out data contamination)
- Mirage mode: asked a visual question with no mention of missing images, the model assumes an image exists, fabricates visual descriptions, and reasons from them. Accuracy is high.
- Guess mode: explicitly told "there is no image, guess the best answer," the model becomes conservative and relies on explicit priors and answer-distribution statistics. Accuracy is low.
- Evaluation: Current multimodal benchmarks are unreliable. ~75% of items test text statistics, not vision.
- Products: Any app with visual input needs missing-modality detection. An API returning 200 doesn't mean the image actually arrived — and the model won't tell you.
- Cognition: Models aren't "understanding images"; they are "navigating the statistical space of visual questions." The image is a switch that triggers navigation, not its fuel.
- Asadi et al. (2026). MIRAGE: The Illusion of Visual Understanding. arXiv:2603.21687. https://arxiv.org/abs/2603.21687
- Related: O'Sullivan et al. (2026). MARCUS: an agentic, multimodal vision-language model for cardiac diagnosis and management. arXiv:2603.22179
- Benchmarks tested: VQA-RAD, MicroVQA, MedXpertQA-MM, MMMU-Pro, ReXVQA
The team tested frontier models — GPT-4.1/5/5.1/5.2, Gemini 2.5/3 Pro, Claude Opus/Sonnet 4.5 — asking visual questions while providing no image at all.
Result: every model "saw" non-existent images. GPT-5.2's mirage rate reached 93.5% — it fabricated detailed visual descriptions in nearly every question.
2. First place with no images
Worse: models without images scored suspiciously high on medical benchmarks.
| Benchmark | Type | No-image accuracy as % of full accuracy | |---|---|---| | VQA-RAD (radiology) | Medical | ~95% | | MicroVQA (microscopy) | Medical | ~85% | | MedXpertQA-MM (general medicine) | Medical | ~90% |
Adding or removing images changes scores by less than 10%. Models primarily rely on non-visual information — question wording, answer-option structure, statistical patterns in training data.
The team trained a text-only model called Super-Guesser:
On a private test set:
| Evaluator | Accuracy | |---|---| | Super-Guesser (3B text-only) | ~75% | | Top multimodal models (average) | ~60-65% | | Radiologists (average) | ~55-60% |
A 3B text-only model beat hundreds-of-billions multimodal models and exceeded human experts by 10 points. It even produced reasoning traces "indistinguishable from genuine visual reasoning" — reviewers couldn't tell.
3. Two modes: Mirage vs. Guess
The same model, same question, differing only in prompt, shows a 10-15% accuracy gap.
GPT-5.1 across three benchmarks:
| Benchmark | Mirage mode | Guess mode | Gap | |---|---|---|---| | MicroVQA | ~55% | ~45% | -10% | | MedXpertQA-MM | ~70% | ~55% | -15% | | MMMU-Pro | ~65% | ~50% | -15% |
Of MMMU-Pro's 23 categories, 18 favored mirage mode. This means the standard "blind-guess control" used in evaluations systematically underestimates benchmark vulnerability.
4. How dangerous are mirages in medicine?
The paper ran 200 repeated tests on Gemini-3-Pro to find its most frequent "phantom" diagnoses:
| Modality | Most common phantom diagnosis | Urgency | |---|---|---| | ECG | Acute STEMI | Immediate intervention | | Dermatology | Malignant melanoma | Urgent excision | | Pathology slides | Various cancers | Cancer treatment |
Consider this scenario: a user uploads a skin photo → a network glitch prevents the upload → the model raises no error and generates phantom reasoning like "the image shows irregular pigmented lesions with unclear borders, consistent with malignant melanoma" → the user receives a detailed diagnosis → unnecessary emergency visits and procedures.
The model doesn't warn you it saw nothing. This is silent failure — users cannot distinguish real analysis from fabrication.
5. B-Clean: scrubbing three-quarters of pseudo-visual questions
The proposed solution, B-Clean: run the benchmark with no images. Any question answerable without an image doesn't test vision — remove it.
| Benchmark | Original items | Compromised | Retained | Removed | |---|---|---|---|---| | MicroVQA | 1,042 | 802 | 240 | 77% | | MedXpertQA-MM | 2,000 | 1,486 | 514 | 74% | | MMMU-Pro | 1,730 | 1,302 | 428 | 75% |
Three-quarters of questions can be answered without images. After cleaning, rankings reshuffled dramatically on MicroVQA:
| Model | Original accuracy | After B-Clean | Drop | |---|---|---|---| | Gemini-3-Pro | 68.8% | 23.2% | -45.6% | | GPT-5.1 | 61.5% | 15.4% | -46.1% |
6. Why this matters
7. The authors' three recommendations
1. Every multimodal evaluation must include no-image controls: systematically disable each input modality, like stress testing. 2. Benchmarks must be kept private: publicly released benchmarks get absorbed into the next generation of pretraining. Rotate items regularly. 3. Measure the with-image vs. without-image gap: absolute accuracy alone is misleading — the gap is the true measure of visual understanding.