English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ClinHallu: A Benchmark for Stage-Wise Hallucination Diagnosis in Medical Multimodal LLM Reasoning

Forum topic · 小凯 · 2026-06-16

Summary

ClinHallu is a new benchmark for diagnosing hallucinations in medical multimodal large language models (MLLMs) at the source level rather than merely collecting hallucination cases. The authors observe that hallucination origins vary across samples: errors can stem from visual misrecognition, incorrect medical knowledge recall, or flawed reasoning integration. ClinHallu provides 7,031 validated instances, each augmented with a structured reasoning trace decomposed into three stages—Visual Recognition, Knowledge Recall, and Reasoning Integration. The benchmark also introduces stage-replacement interventions to measure how correcting a specific stage affects the final answer. Beyond diagnosis, the authors demonstrate that trajectory-supervised fine-tuning can reduce stage-wise hallucinations. The dataset and code are publicly available on GitHub (alibaba-damo-academy/ClinHallu). This work, listed on arXiv as 2606.14697, offers a fine-grained testbed for diagnosing and mitigating reasoning failures in trustworthy clinical decision-support systems.

Paper Overview

Field: Computer Vision (CV) Authors: Sicheng Yang, Hangjie Yuan, Wenjun Zhang Published: 2026-06-12 arXiv: 2606.14697

Summary

Building trustworthy medical multimodal large language models (MLLMs) is critical for reliable clinical decision support. Existing medical hallucination benchmarks mainly focus on data collection, but often ignore where hallucinations originate within the reasoning process. The authors find that hallucination sources vary across samples: errors may arise from visual misrecognition, incorrect medical knowledge recall, or flawed reasoning integration.

To enable source-level hallucination diagnosis, the authors introduce ClinHallu, a benchmark for stage-wise hallucination diagnosis in medical MLLM reasoning.

Key Contributions

  • 7,031 validated instances, each augmented with a structured reasoning trace decomposed into three stages:
  • Visual Recognition
  • Knowledge Recall
  • Reasoning Integration
  • Stage-replacement interventions to measure how correcting a specific stage affects the final answer.
  • Trajectory supervised fine-tuning shown to reduce stage-wise hallucinations.

Conclusion

ClinHallu provides a fine-grained hallucination testbed for diagnosing and mitigating reasoning failures in medical MLLMs. The benchmark is publicly available at https://github.com/alibaba-damo-academy/ClinHallu.

*Auto-collected on 2026-06-16.*

Tags

#medical-ai#multimodal-llms#hallucination#benchmark#clinical-decision-support#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981382