English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

HerHealthEval: A Multilingual, Register-Sensitive Benchmark for Women's Health Communication Understanding in LLMs

Forum topic · 小凯 · 2026-09-20

Summary

HerHealthEval is a controlled evaluation framework for assessing how well large language models understand women's health communication across languages and registers. Each clinical case is provided in matched English, French, and Modern Standard Arabic versions using six communicative forms: canonical, clinical, layperson, indirect/hedged, emotionally concerned, and deliberately under-specified. The first five preserve identical clinical information, while the under-specified form intentionally omits details to test whether models recognize the need for clarification. The paper evaluates a multilingual instruction-tuned model and its QLoRA-adapted variants on concern classification, risk calibration, clarification behavior, parsing compliance, and cross-form consistency. Results show that aggregate accuracy can mask safety-critical failures: one multilingual-adapted model exhibited an under-triage rate of 0.994 in French and Arabic; controlled re-adaptation using language-invariant, source-language risk labels reduced under-triage to 0.572 and 0.558 respectively. The findings argue that robust multilingual medical evaluation must explicitly test register variation, uncertainty handling, and the provenance and invariance of adaptation labels. arXiv: 2609.20684.

Paper Overview

  • Field: NLP
  • Authors: Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi, Hana Essam Sayed Ahmed Amrya, Mariam Mousa
  • Published: 2026-09-17
  • arXiv: 2609.20684
  • Summary

    Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. The authors introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication.

    For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms:

    1. Canonical 2. Clinical 3. Layperson 4. Indirect or hedged 5. Emotionally concerned 6. Deliberately under-specified

    The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed.

    The framework evaluates a multilingual instruction-tuned model and its QLoRA-adapted variants on:

  • Concern classification
  • Risk calibration
  • Clarification behavior
  • Parsing compliance
  • Cross-form consistency
  • Key Findings

  • Aggregate accuracy and consistency can mask safety-critical failures: one multilingual-adapted model exhibited an under-triage rate of 0.994 in French and Arabic.
  • Controlled re-adaptation using language-invariant risk labels derived from the source language reduced under-triage to 0.572 (French) and 0.558 (Arabic).
  • Robust multilingual medical evaluation must explicitly test register variation, uncertainty handling, and the provenance and invariance of adaptation labels.
---

*Auto-collected on 2026-09-20.*

Tags

#nlp#large-language-models#healthcare#multilingual#evaluation-benchmark#women-health#qlora#model-safety

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635019