English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MRI-Eval: A Tiered Benchmark for LLM Performance on MRI Physics and GE Scanner Operations

Forum topic · 小凯 · 2026-05-08

Summary

MRI-Eval (arXiv:2605.05175) is a tiered benchmark developed by Perry E. Radau to evaluate large language models on MRI physics and GE scanner operations knowledge. It contains 1,365 scored multiple-choice items across nine categories and three difficulty tiers, drawn from textbooks, GE scanner manuals, programming course materials, and expert-generated questions. Five model families were tested: GPT-5.4, Claude Opus 4.6, Claude Sonnet 4.6, Gemini 2.5 Pro, and Llama 3.3 70B. Under standard MCQ conditions, overall accuracy ranged from 93.2% to 97.1%, with GE scanner operations the weakest category for every model (88.2% to 94.6%). In stem-only (free-text) conditions, frontier-model accuracy dropped to 58.4% to 61.1% and Llama 3.3 70B to 37.1%, while GE scanner operations stem-only accuracy fell to just 13.8% to 29.8%. The findings show that strong MCQ scores can mask weak free-text recall, especially for vendor-specific operational knowledge, and caution against using raw LLM outputs for GE-specific protocol guidance.

MRI-Eval (arXiv: 2605.05175) by Perry E. Radau is a tiered benchmark for relatively comparing LLM performance on MRI physics and GE scanner operations knowledge.

Background

Existing MRI LLM benchmarks rely mainly on review-book multiple-choice questions (MCQs), on which top proprietary models already score highly, limiting discrimination. No systematic benchmark had evaluated vendor-specific scanner operational knowledge central to research MRI practice.

Methods

  • 1,365 scored MCQ items across nine categories and three difficulty tiers
  • Sources: textbooks, GE scanner manuals, programming course materials, and expert-generated questions
  • Five model families evaluated: GPT-5.4, Claude Opus 4.6, Claude Sonnet 4.6, Gemini 2.5 Pro, Llama 3.3 70B
  • MCQ was the primary condition; a stem-only condition removed options and used an independent LLM judge; a primed stem-only condition tested responses to incorrect user claims
  • Results

  • Overall MCQ accuracy: 93.2% to 97.1%
  • GE scanner operations was the lowest-scoring category for every model (88.2% to 94.6%)
  • In stem-only conditions, frontier-model accuracy fell to 58.4% to 61.1%; Llama 3.3 70B fell to 37.1%
  • GE scanner operations stem-only accuracy was only 13.8% to 29.8%

Conclusion

High MCQ performance can mask weak free-text recall, especially for vendor-specific operational knowledge. MRI-Eval is most informative as a relative comparison benchmark rather than an absolute competency measure, and the results support caution when using raw LLM outputs for GE-specific protocol guidance.

Tags

#llm-benchmark#mri-physics#medical-imaging#ge-scanner#model-evaluation#arxiv#machine-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619590