English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DeepSciVerify: Verifying Scientific Claim-Citation Alignment with Two-Stage Selective Evidence Escalation

Forum topic · 小凯 · 2026-05-29

Summary

DeepSciVerify is a two-stage pipeline for verifying whether scientific claims are supported by their cited evidence, a common failure mode in reports generated by large language models. The system first performs abstract-level claim verification and defers judgment on uncertain cases, escalating only those to full-text passage-level retrieval and analysis. This selective escalation leverages complementary behaviors across LLMs, since some models are more conservative while others are more decisive under uncertainty. On the SCitance benchmark, DeepSciVerify achieves 86.7 Micro-F1, outperforming strong abstract-only baselines by +4.5 points while resolving 67% of instances without full-text retrieval. The results indicate that selective evidence escalation improves both the accuracy and efficiency of claim-citation verification, enhancing the reliability of LLM-generated scientific reports in high-stakes settings. The paper (arXiv:2605.27710) is authored by Shaghayegh Sadeghi, Khashayar Khajavi, Rise Adhikari, et al. in the NLP field.

Paper Overview

  • Field: NLP
  • Authors: Shaghayegh Sadeghi, Khashayar Khajavi, Rise Adhikari, et al.
  • Published: 2026-05-28
  • arXiv: 2605.27710
  • Summary

    Misalignment between claims and their cited evidence is a common failure mode in reports generated by large language models, limiting their reliability in scientific and other high-stakes settings. This paper presents DeepSciVerify, a two-stage pipeline for scientific claim-citation verification that combines abstract-level reasoning with selective escalation to passage-level evidence.

    The system first verifies claims using the paper abstract and defers uncertain cases, retrieving and analyzing full-text passages only when necessary. This design leverages complementary behaviors across LLMs, as some models are more conservative while others are more decisive under uncertainty.

    Key Results

  • On the SCitance benchmark, DeepSciVerify achieves 86.7 Micro-F1.
  • This outperforms strong abstract-only baselines by +4.5 points.
  • 67% of instances are resolved without any full-text retrieval, improving efficiency.

Conclusion

The results suggest that selective evidence escalation improves both the accuracy and the efficiency of claim-citation verification, pointing toward more reliable LLM-assisted scientific writing and review workflows.

Tags

#nlp#llm#claim-verification#citation-alignment#scientific-text-processing#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980519