Paper Overview
Field: NLP Authors: Mariano Barone, Francesco Di Serio, Roberto Moio Published: 2026-04-22 arXiv: 2604.20791
Abstract
Large Language Models (LLMs) are increasingly deployed in healthcare, yet their communicative alignment with clinical standards remains insufficiently quantified. The authors conduct a multidimensional evaluation of general-purpose and domain-specialized LLMs across structured medical explanations and real-world physician-patient interactions, analyzing semantic fidelity, readability, and affective resonance.
Key Findings
- Emotional amplification: Baseline models amplify affective polarity relative to physicians (Very Negative: 43.14–45.10% vs. 37.25%).
- Higher linguistic complexity: Larger architectures such as GPT-5 and Claude produce substantially higher linguistic complexity (FKGL up to 16.91–17.60 vs. 11.47–12.50 in physician-authored responses).
- Empathy-oriented prompting: Reduces extreme negativity and lowers grade-level complexity (up to -6.87 FKGL points for GPT-5), but does not significantly increase semantic fidelity.
- Collaborative rewriting: Produces the strongest overall alignment. Rewriting configurations achieve the highest semantic similarity to physician answers (average up to 0.93) while consistently improving readability and reducing affective extremity.
- Dual-stakeholder evaluation: No model surpasses physicians on cognitive criteria, while patients consistently prefer rewritten variants for clarity and emotional tone.
Conclusion
The findings suggest that LLMs function most effectively as collaborative communication enhancers rather than replacements for clinical expertise.
---
*Auto-collected on 2026-04-24*