Summary
MRI-Eval (arXiv:2605.05175) is a tiered benchmark developed by Perry E. Radau to evaluate large language models on MRI physics and GE scanner operations knowledge. It contains 1,365 scored multiple-choice items across nine categories and three difficulty tiers, drawn from textbooks, GE scanner manuals, programming course materials, and expert-generated questions. Five model families were tested: GPT-5.4, Claude Opus 4.6, Claude Sonnet 4.6, Gemini 2.5 Pro, and Llama 3.3 70B. Under standard MCQ conditions, overall accuracy ranged from 93.2% to 97.1%, with GE scanner operations the weakest category for every model (88.2% to 94.6%). In stem-only (free-text) conditions, frontier-model accuracy dropped to 58.4% to 61.1% and Llama 3.3 70B to 37.1%, while GE scanner operations stem-only accuracy fell to just 13.8% to 29.8%. The findings show that strong MCQ scores can mask weak free-text recall, especially for vendor-specific operational knowledge, and caution against using raw LLM outputs for GE-specific protocol guidance.
MRI-Eval (arXiv: 2605.05175) by Perry E. Radau is a tiered benchmark for relatively comparing LLM performance on MRI physics and GE scanner operations knowledge.
Background
Existing MRI LLM benchmarks rely mainly on review-book multiple-choice questions (MCQs), on which top proprietary models already score highly, limiting discrimination. No systematic benchmark had evaluated vendor-specific scanner operational knowledge central to research MRI practice.
Methods
- 1,365 scored MCQ items across nine categories and three difficulty tiers
- Sources: textbooks, GE scanner manuals, programming course materials, and expert-generated questions
- Five model families evaluated: GPT-5.4, Claude Opus 4.6, Claude Sonnet 4.6, Gemini 2.5 Pro, Llama 3.3 70B
- MCQ was the primary condition; a stem-only condition removed options and used an independent LLM judge; a primed stem-only condition tested responses to incorrect user claims
Results
- Overall MCQ accuracy: 93.2% to 97.1%
- GE scanner operations was the lowest-scoring category for every model (88.2% to 94.6%)
- In stem-only conditions, frontier-model accuracy fell to 58.4% to 61.1%; Llama 3.3 70B fell to 37.1%
- GE scanner operations stem-only accuracy was only 13.8% to 29.8%
Conclusion
High MCQ performance can mask weak free-text recall, especially for vendor-specific operational knowledge. MRI-Eval is most informative as a relative comparison benchmark rather than an absolute competency measure, and the results support caution when using raw LLM outputs for GE-specific protocol guidance.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177619590