English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Operadic Consistency: A Label-Free Signal for LLM Compositional Reasoning

Forum topic · 小凯 · 2026-06-14

Summary

A new paper (arXiv:2606.13649) by Nathaniel Bottman, Yinhong Liu, and Kyle Richardson introduces operadic consistency (OC), a label-free diagnostic for detecting LLM reasoning failures. Drawing on operad theory, the method checks whether a model's direct answer to a compositional query agrees with the answer produced by composing a stated decomposition of that query. Unlike confidence-based baselines, OC requires no ground-truth labels. Across 12 LLMs evaluated on 4 multi-hop QA datasets, OC scores correlate strongly with accuracy, with Pearson r between 0.86 and 0.94. The authors also show that OC improves selective prediction, outperforming tuned confidence baselines. This makes OC a practical, unsupervised signal for assessing and screening compositional reasoning reliability in large language models.

Paper Overview

  • Field: NLP
  • Authors: Nathaniel Bottman, Yinhong Liu, Kyle Richardson
  • Posted: 2026-06-11
  • arXiv: 2606.13649
  • Abstract

    Detecting LLM reasoning failures without ground-truth labels has motivated confidence baselines. Operad theory suggests a complementary diagnostic: a model's direct answer should agree with its answer by composing a stated decomposition. We instantiate this as operadic consistency (OC), strongly correlated with accuracy (Pearson r in [0.86, 0.94]) across 12 LLMs on 4 multi-hop QA datasets. OC yields selective-prediction improvements over tuned baselines.

    Key Points

  • Introduces operadic consistency (OC), a label-free diagnostic inspired by operad theory.
  • Core idea: a model's direct answer to a compositional query should agree with the answer obtained by composing its own stated decomposition of the query.
  • Evaluated on 12 LLMs across 4 multi-hop QA datasets; OC correlates strongly with accuracy (Pearson r ∈ [0.86, 0.94]).
  • OC improves selective prediction and outperforms tuned confidence-based baselines.
---

*Auto-collected on 2026-06-14.*

Tags

#llm#reasoning#nlp#multi-hop-qa#evaluation#arxiv#papers

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981286