English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Verifiable by Construction: Claim-Level Evaluation of Verbatim Citations in Clinical QA with LLMs

Forum topic · 小凯 · 2026-09-16

Summary

Large language models are increasingly used for clinical question answering, but typical citations point to broad reference texts that time-pressed clinicians cannot efficiently verify. This paper from Columbia-affiliated researchers (Jiashuo Zhang, Yuling Chen, Yvonne Commodore-Mensah, Michael Oberst, arXiv:2609.15964) evaluates whether current LLMs can produce answers that are verifiable by construction, meaning every factual claim is supported by fine-grained verbatim quotes from reference material. The authors build a standardized evaluation harness based on four clinical practice guidelines and test 12 LLMs on 222 synthetic clinical questions, measuring each stage: citing every claim, generating verbatim quotes, and ensuring quotes fully substantiate claims. Results show most models, except lightweight ones like claude-haiku-4.5, can attach verbatim quotes to over 90% of claims via prompting alone. However, quotes often fail to fully substantiate all details of the attached claim: claude-opus-5 produced verbatim quotes for 98.0% of claims, but only 37.1% of those quotes fully substantiated their claims. The work quantifies capability gaps in building verifiable clinical QA systems and releases artifacts for future research.

Overview

Field: NLP Authors: Jiashuo Zhang, Yuling Chen, Yvonne Commodore-Mensah, Michael Oberst Published: 2026-09-14 arXiv: 2609.15964

Abstract

Large language models (LLMs) have been widely adopted for clinical question answering (QA). Current systems can attach citations to their answers, but these often point to broad texts, leaving time-pressed clinicians unable to verify them efficiently. An alternative is to ensure that responses are verifiable by construction: providing fine-grained verbatim quotes from reference material that substantiate claims, so users can verify an answer without opening other documents.

In this paper, the authors evaluate the ability of current models to perform this task end-to-end: from providing citations for every factual claim, to producing verbatim quotes, to ensuring that those quotes fully substantiate the claims.

Methodology

  • Built a standardized evaluation harness over four clinical practice guidelines
  • Evaluated 12 LLMs on 222 synthetic clinical questions
  • Measured each stage separately: claim citation coverage, quote verbatimness, and full substantiation of claims by their quotes
  • Key Findings

  • Apart from lightweight models such as claude-haiku-4.5, most models can attach verbatim quotes to over 90% of claims with prompting alone.
  • However, quotes often fail to substantiate all details of the attached claim. For example, claude-opus-5 generated verbatim quotes for 98.0% of claims, but only 37.1% of claims were fully substantiated by their quotes.
  • The paper provides insights into the capability gaps of current LLMs for building verifiable clinical QA systems, along with artifacts for future research.
---

*Auto-collected on 2026-09-16*

Tags

#nlp#large-language-models#clinical-qa#citations#evaluations#hallucination#verifiability#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634868