English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Two Independent Axes of LLM Abstention: Why a Single Confidence Threshold Fails

Forum topic · ✨步子哥 · 2026-07-10

Summary

A forum post discusses Benedikt Wagner's July 2026 paper 'Two Axes of LLM Abstention,' which argues that answer correctness and question answerability are orthogonal dimensions that current selective-answering mechanisms conflate. Experiments on five instruction-tuned models (2B–14B, three families) show output confidence barely predicts answerability (AUROC 0.54–0.67, near random), while linear probes on hidden states reach 0.97–0.99 — with no scaling trend. On the CREPE dataset of real user questions with false premises, model self-evaluation methods achieve at best 0.67 AUROC. Prompting models to check premises causes false challenges on 57% of legitimate questions; routing premise checks through hidden-state probes triples challenge precision. The paper formalizes abstention as three-class selective acceptance with separate budgets for unanswerable (α_U) and wrong-answerable (α_W) questions, certified via binomial bounds. Results reveal an asymmetric scaling effect: abstention on unanswerable questions is detection-limited, while abstention on wrong answers is capability-limited.

You ask an LLM: "Who is the current King of France?"

It confidently answers: "Louis XIX."

Wrong — but not the kind of "wrong answer" wrong. It's the kind of "this question should never have been answered" wrong. France has no king. The question itself rests on a false premise.

Now ask: "What is 2+2?"

It answers: "5."

Also wrong — but this time the question is legitimate and only the answer is incorrect.

Both look like "the model gave a wrong response," but they are fundamentally different. One is an answer problem, the other is a question problem. Benedikt Wagner (City St George's, University of London), in his July 2026 paper *Two Axes of LLM Abstention*, proves that these two error types are orthogonal axes, and existing abstention mechanisms conflate them.

One Ruler Measuring Two Things

The dominant selective-answering approach is simple: set a confidence threshold; answer below it is abstained. Like a thermometer that triggers an alarm above a certain temperature.

But Wagner found that this "thermometer" measures two different things:

  • Answer correctness: Can I answer this correctly?
  • Question answerability: Should this question be answered at all?
  • The paper defines three question classes:

  • C (correct-answerable): answerable and answered correctly
  • W (wrong-answerable): answerable but answered incorrectly
  • U (unanswerable): not answerable (false premises, no solution, etc.)
  • Ideally, models should abstain on both W and U. But across 5 instruction-tuned models (2B–14B, from 3 model families), Wagner found a disturbing fact:

    The models' output confidence is nearly blind to question answerability.

    Specifically:

  • Output confidence predicting answerability: AUROC 0.54–0.67 (close to random guessing)
  • Linear probes on hidden states predicting answerability: AUROC 0.97–0.99
  • Scaling trend: none. The 14B model is no better than the 2B at distinguishing answerable from unanswerable questions.
  • Stranger still: even a trained direction that separates C from W (correct vs. wrong answers) barely helps distinguish answerable from unanswerable questions. "Answer correctness" and "question answerability" are entangled in the model's representation space — pull out a direction to separate them and you find they don't even lie on the same line.

    The CREPE Dataset: A Test with Real User Questions

    You might think these "unanswerable" questions are artificially constructed — of course the model hasn't seen them. Wagner's use of the CREPE dataset dispels this.

    CREPE questions come from real users, some containing false premises. For example: "Who is the author of the novel 'XXX' that topped the New York Times bestseller list?" — if that novel never topped the list, the question is unanswerable.

    On CREPE:

  • Trained output readouts: nearly random
  • Model self-evaluations (P(IK), P(True), direct premise checks): at best 0.67 AUROC
  • Internal linear readouts (hidden-state probes): 0.69–0.77
  • What does 0.67 mean? A coin flip gives you 0.5. The model's own judgment of "should I answer this" is only slightly better than chance.

    A 57% False-Challenge Rate: Good Intentions Gone Wrong

    So what happens if you just make the model "check premises" before answering?

    Wagner ran this experiment. The result:

    The model challenged 57% of perfectly legitimate questions.

    Unable to tell which questions have false premises, it simply suspected all of them. Like an oversensitive smoke alarm — it goes off for both cooking and fires, until you rip it out.

    But Wagner found a fix: route with hidden-state linear probes. Let the probe first judge "is this answerable," and only trigger premise-checking when the probe says "unanswerable." The result:

    Challenge precision improved 3×.

    In other words, the answerability signal exists in the hidden states (AUROC 0.97–0.99) — the model just can't use it on its own. You need to attach a probe from the outside to read it.

    Dual-Budget Certification: Turning Abstention into an Engineering Problem

    The paper's most elegant move is formalizing abstention as three-class selective acceptance with two independent budgets:

  • α_U = 0.15: at most 15% of unanswerable questions are answered erroneously
  • α_W = 0.50: at most 50% of answerable questions are answered incorrectly
  • δ = 0.10: a statistical guarantee at 90% confidence
  • Certified with exact binomial bounds and validated on a held-out test set independent of training and threshold tuning.

    The beauty of the framework: it acknowledges the two error types have different costs and should be controlled separately. Like factory quality control — you don't lump "defective product function" and "wrong shipment" into a single defect rate, because their causes and fixes differ entirely.

    An Asymmetric Scaling Effect

    The certification results expose a "scale-dependent asymmetry":

  • Unanswerable budget: 4 of 5 models passed certification
  • Wrong-answer budget: bottlenecked by the models' intrinsic accuracy
On the 8B model, a factorized dual-threshold strategy passed both budgets simultaneously. But on smaller models, the wrong-answer budget is capped by model capability — no better abstention strategy can compensate for a model that simply isn't strong enough.

Practical implication: for small models, the abstention ceiling is capability; for large models, the abstention ceiling is detection. Past a certain scale, the bottleneck shifts from "can the model answer correctly" to "does it know whether it should answer."

The Paper's Deeper Contribution

This paper does something rare: it shows that a widely used framework (single-threshold confidence abstention) is conceptually wrong — not fixable by tuning.

It's not "the threshold isn't good enough." It's that "using one ruler to measure two orthogonal things" is the flawed idea itself.

This echoes the Co-Failure Ceiling paper's argument — not "the composition strategy isn't good enough," but "using ρ to measure diversity" is the wrong metric. Both papers say the same thing: when the conceptual dimensions are wrong, no amount of fine-grained engineering optimization can save you.

Wagner's factorized framework turns abstention from a vague "when to stay silent" question into two clear engineering problems: how to control the misanswer rate on unanswerable questions, and how to control the error rate on answerable ones. The former relies on detection (hidden-state probes); the latter on capability (the model itself). The two can be optimized and certified independently.

It's far more boring than "tune a confidence threshold" — but it maps onto the model's true failure structure. Good engineering is often like this: get the concepts right first, then optimize.

---

Paper: Two Axes of LLM Abstention: Answer Correctness and Question Answerability Author: Benedikt J. Wagner (City St George's, University of London) Date: July 9, 2026

Tags

#llm#abstention#selective-prediction#false-premises#confidence-calibration#model-alignment#ai-safety#measurement-validity

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346298