[论文] Verifiable by Construction: Claim-Level Evaluation of Verbatim Citatio...
研究领域: NLP 作者: Jiashuo Zhang, Yuling Chen, Yvonne Commodore-Mensah, Michael Oberst 发布时间: 2026-09-14 arXiv: 2609.15964
论文概要
研究领域: NLP 作者: Jiashuo Zhang, Yuling Chen, Yvonne Commodore-Mensah, Michael Oberst 发布时间: 2026-09-14 arXiv: 2609.15964
中文摘要
大语言模型已被广泛用于临床问答。当前系统可以为答案附加引用,但这些引用往往指向宽泛的文本,使时间紧迫的临床医生无法高效验证。一种替代方案是确保回应"构造即可验证":提供来自参考材料的细粒度原文引用,使用户无需打开其他文档即可验证答案。本文端到端评估了当前模型执行此任务的能力:从为每个事实性声明提供引用,到生成原文引述,再到确保这些引述完整支撑声明。为此,我们基于四份临床实践指南构建了标准化评估框架,在222个合成临床问题上评估了12个LLM,分别测量上述各阶段。我们发现,除 claude-haiku-4.5 等轻量模型外,大多数模型仅凭提示即可为超过90%的声明附加原文引述。然而,这些引述往往无法完整支撑所附声明的所有细节。例如,claude-opus-5 为98.0%的声明生成了原文引述,但完全支撑的仅占37.1%。我们的工作提供了关于当前LLM在构建可验证临床QA系统方面能力缺口的洞察,并提供了供未来研究的工件。
原文摘要
Large language models (LLMs) have been widely adopted for clinical question answering (QA). Current systems can attach citations to their answers, but these often point to broad texts, leaving time-pressed clinicians unable to verify them efficiently. An alternative is to ensure that responses are verifiable by construction: providing fine-grained verbatim quotes from reference material that substantiate claims, so users can verify an answer without opening other documents. In this paper, we evaluate the ability of current models to perform this task end-to-end: from providing citations for every factual claim, to producing verbatim quotes, to ensuring that those quotes fully substantiate the claims. To do so, we build a standardized harness over four clinical practice guidelines and evalu...
*自动采集于 2026-09-16*
#论文 #arXiv #NLP #小凯