[论文] HerHealthEval: Evaluating Multilingual and Register-Sensitive Understa...
研究领域: NLP 作者: Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi, Hana Essam Sayed Ahmed Amrya, Mariam Mousa 发布时间: 2026-09-17 arXiv: 2609.2…
论文概要
研究领域: NLP 作者: Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi, Hana Essam Sayed Ahmed Amrya, Mariam Mousa 发布时间: 2026-09-17 arXiv: 2609.20684
中文摘要
大语言模型日益用于医疗沟通,但多数评估关注回答质量,默认用户的诉求已被正确理解。我们提出 HerHealthEval——一个面向女性健康沟通多语言理解的受控评估框架。每个临床病例提供英语、法语与现代标准阿拉伯语的匹配版本,使用六种交际形式:规范型、临床型、外行型、间接/委婉型、情绪担忧型与故意欠说明型。前五种表达同一潜在诉求、保留相同临床信息;欠说明型有意省略相关信息,测试模型能否识别需要澄清。我们在诉求分类、风险校准、澄清行为、解析合规性与跨形式一致性上评估了一个多语言指令模型及其 QLoRA 适配变体。结果揭示:聚合准确率与一致性会掩盖安全攸关的失效——某多语言适配模型在法语和阿拉伯语下出现 0.994 的分诊不足;使用源自源语、语言不变的风险标签进行受控重适配后,分诊不足分别降至 0.572 与 0.558。这些发现表明:稳健的多语言医疗评估必须显式测试语域变异、不确定性处理,以及适配标签的来源与不变性。
原文摘要
Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms: canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified. The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed. W...
*自动采集于 2026-09-20*
#论文 #arXiv #NLP #小凯