English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Diagnosing LLM Judge Reliability via Conformal Prediction Sets and Transitivity Analysis

Forum topic · 小凯 · 2026-04-18

Summary

This paper investigates the per-instance reliability of LLM-as-judge frameworks used for automatic natural language generation evaluation. Using SummEval as a benchmark, the authors propose a two-part diagnostic toolkit. First, transitivity analysis reveals that despite low aggregate violation rates of 0.8%-4.1%, 33%-67% of documents exhibit at least one directed 3-cycle, indicating widespread inconsistency. Second, split conformal prediction sets over 1-5 Likert scores provide theoretically guaranteed coverage of at least (1-alpha), with set width functioning as a per-instance reliability indicator (Spearman rs = +0.576, N=1,918, p < 10^-100 across judges). Set width shows consistent cross-judge agreement (mean r = 0.32-0.38), showing it reflects document-level difficulty. Across four judges and four criteria, criteria matter more than judges: relevance is most reliably judged (mean set size about 3.0), coherence is next (about 3.9), while fluency and consistency remain unreliable (about 4.9). Code, prompts, and cached results are released.

Paper Overview

  • Field: NLP
  • Authors: Manan Gupta, Dhruv Kumar
  • Published: 2025-04-17
  • arXiv: 2504.13084
  • Abstract

    LLM-as-judge frameworks are increasingly used for automatic NLG evaluation, yet their per-instance reliability remains poorly understood. We present a two-pronged diagnostic toolkit applied to SummEval:

    1. Transitivity analysis that reveals widespread per-input inconsistency masked by low aggregate violation rates (ρ̄ = 0.8%-4.1%), with 33%-67% of documents exhibiting at least one directed 3-cycle. 2. Split conformal prediction sets over 1-5 Likert scores providing theoretically-guaranteed ≥(1-α) coverage, with set width serving as a per-instance reliability indicator (rs = +0.576, N=1,918, p < 10^-100, pooled across all judges).

    Critically, prediction set width shows consistent cross-judge agreement (r̄ = 0.32-0.38), demonstrating it captures document-level difficulty rather than judge-specific noise.

    Key Findings

  • Criteria matter more than judges: Across four judges and four criteria, both diagnostics converge: criteria exert stronger influence than judge choice.
  • Per-criterion reliability ranking:
  • Relevance: most reliably judged (mean set size ≈ 3.0)
  • Coherence: second (mean set size ≈ 3.9)
  • Fluency and consistency: remain unreliable (mean set size ≈ 4.9)
  • Author contributions: The authors release all code, prompts, and cached results to support reproducibility.
  • Resources

  • arXiv link: https://arxiv.org/abs/2504.13084
*Auto-collected 2026-04-18*

#paper #arXiv #NLP

Tags

#llm-as-judge#conformal-prediction#transitivity-analysis#nlg-evaluation#summeval#reliability-diagnostics#nlp-evaluation#arxiv-2504.13084

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618541