Paper Overview
- Field: NLP
- Authors: Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi, Hana Essam Sayed Ahmed Amrya, Mariam Mousa
- Published: 2026-09-17
- arXiv: 2609.20684
- Concern classification
- Risk calibration
- Clarification behavior
- Parsing compliance
- Cross-form consistency
- Aggregate accuracy and consistency can mask safety-critical failures: one multilingual-adapted model exhibited an under-triage rate of 0.994 in French and Arabic.
- Controlled re-adaptation using language-invariant risk labels derived from the source language reduced under-triage to 0.572 (French) and 0.558 (Arabic).
- Robust multilingual medical evaluation must explicitly test register variation, uncertainty handling, and the provenance and invariance of adaptation labels.
Summary
Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. The authors introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication.
For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms:
1. Canonical 2. Clinical 3. Layperson 4. Indirect or hedged 5. Emotionally concerned 6. Deliberately under-specified
The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed.
The framework evaluates a multilingual instruction-tuned model and its QLoRA-adapted variants on:
Key Findings
*Auto-collected on 2026-09-20.*