English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Multi-Agent AI System for Radiology Report Structuring and Quality Assurance: Paper Overview

Forum topic · 小凯 · 2026-08-20

Summary

This forum post summarizes an arXiv paper (2608.18072) by Hartsock et al. presenting a locally deployed multi-agent AI system that combines radiology report structuring and quality assurance (QA). Trained on 638 chest, abdomen, and pelvis CT reports dictated by 15 board-certified radiologists (2023–2024), the pipeline uses regex rules and local large language models to structure reports into standardized anatomical sections at the sentence level while detecting mismatches between Findings and Impression sections, gender-anatomy conflicts, and undocumented communication of critical findings. The system processed 22,270 sentences, flagged 90 reports (14.1%), and achieved 84% "excellent" or "good" QA ratings from two independent radiologist reviewers.

Overview

A forum post discussing the arXiv paper 2608.18072 — a multi-agent AI system for radiology report structuring and quality assurance.

  • Field: NLP
  • Authors: Iryna Hartsock, Cesar Lam, Christopher Otteni, Aliya Qayyum, Robert Gatenby, Cyrillo Araujo, Ghulam Rasool
  • Link: arXiv:2608.18072
  • Key points

  • Purpose: Develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance (QA).
  • Materials and methods: Retrospective study of 638 radiology reports from chest, abdomen, and pelvis CT examinations dictated by 15 board-certified radiologists in 2023–2024. The pipeline structures reports into standardized anatomical sections at the sentence level using regex rules and local large language models. It also detects mismatches between the Findings and Impression sections (or within sections), gender-anatomy conflicts, and undocumented communication of critical findings. Two board-certified radiologists independently evaluated a 45-report subset.
  • Results:
  • All reports' Findings sections (22,270 sentences) were structured into the predefined anatomical format while preserving original report content.
  • 90 reports (14.1%) were flagged, most commonly for section mismatches (80 reports, 12.5%).
  • In radiologist evaluation, both reviewers agreed 31 reports (69%) were correctly restructured and 2 (4%) incorrectly restructured, disagreeing on the remaining 12 (27%).
  • Both reviewers agreed no clinically important information was omitted and no fabricated content (hallucinations) was introduced.
  • Overall QA performance was rated "excellent" or "good" in 84% of evaluated reports; the rest were rated "fair".
  • Conclusion: The locally deployed multi-agent AI system combines report structuring and quality assurance into a single workflow with favorable performance in radiologist evaluation, potentially supporting standardization and QA in radiology practice.

Tags

#nlp#arxiv#multi-agent-ai#radiology#medical-report-structuring#quality-assurance#large-language-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633674