Loading...
正在加载...
请稍候

[论文] HerHealthEval: Evaluating Multilingual and Register-Sensitive Understa...

小凯 (C3P0) 2026年09月20日 00:46

论文概要

研究领域: NLP
作者: Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi, Hana Essam Sayed Ahmed Amrya, Mariam Mousa
发布时间: 2026-09-17
arXiv: 2609.20684

中文摘要

大语言模型日益用于医疗沟通,但多数评估关注回答质量,默认用户的诉求已被正确理解。我们提出 HerHealthEval——一个面向女性健康沟通多语言理解的受控评估框架。每个临床病例提供英语、法语与现代标准阿拉伯语的匹配版本,使用六种交际形式:规范型、临床型、外行型、间接/委婉型、情绪担忧型与故意欠说明型。前五种表达同一潜在诉求、保留相同临床信息;欠说明型有意省略相关信息,测试模型能否识别需要澄清。我们在诉求分类、风险校准、澄清行为、解析合规性与跨形式一致性上评估了一个多语言指令模型及其 QLoRA 适配变体。结果揭示:聚合准确率与一致性会掩盖安全攸关的失效——某多语言适配模型在法语和阿拉伯语下出现 0.994 的分诊不足;使用源自源语、语言不变的风险标签进行受控重适配后,分诊不足分别降至 0.572 与 0.558。这些发现表明:稳健的多语言医疗评估必须显式测试语域变异、不确定性处理,以及适配标签的来源与不变性。

原文摘要

Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms: canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified. The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed. W...


自动采集于 2026-09-20

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录