When Quantization Preserves Accuracy but Not Evidence: The "Proxy Objective Trap" in Medical LLMs
A Clinical Scenario
Imagine a doctor using an LLM to assist diagnosis. The patient is a 60-year-old male with chest pain and sweating; the ECG shows ST-segment elevation. The model answers "acute myocardial infarction" with the rationale: "ST-segment elevation indicates transmural ischemia, consistent with acute myocardial infarction."
The doctor nods — the answer is right, the reasoning is right, it's trustworthy.
Now the model is quantized for deployment on edge devices, with weights compressed from 16-bit to 4-bit. The quantized model still answers "acute myocardial infarction." Accuracy unchanged. But the rationale changes — it now says: "The patient is over 50, chest pain is a common symptom, consider acute myocardial infarction."
The answer is still correct. But the reasoning has shifted from "ST-segment elevation, the key clinical evidence" to generic observations like "age and chest pain." If the doctor trusts this rationale to understand the model's decision basis, they might conclude that "age + chest pain" alone suffices to diagnose a heart attack — clinically problematic.
This is the problem revealed by the paper "When Quantization Preserves Accuracy but Not Evidence": quantization preserves the model's answers while quietly discarding the evidence the model relied on to produce them.
Core Finding: Accuracy and Rationale Support Decouple
The paper's core diagnostic experiment can be summarized in one table:
| Token Type | Share in Rationale | Share in Quantization Loss (KL Divergence) | |---|---|---| | Generic/discourse words | 75.3% | 83.1% | | Clinical cue words | 22.2% | 15.6% | | Numerical cues | 0.67% | 0.38% | | Answer/option words | 1.80% | 0.99% |
How to read this: the distribution of quantization loss across token types does not match the tokens' actual distribution. Generic discourse words ("therefore," "in summary," "based on the above") make up 75.3% of rationales but account for 83.1% of quantization loss — they are over-protected. Clinical cue words ("ST-segment elevation," "troponin," "cardiogenic shock") make up 22.2% of rationales but only 15.6% of quantization loss — they are relatively neglected.
What does this mean? The standard quantization objective (minimizing KL divergence) mainly protects tokens that "make sentences fluent," not tokens that "make rationales credible." The quantized model's rationales look smooth and plausible, but the key clinical evidence supporting the answer is diluted.
This is what the title means by "Preserves Accuracy but Not Evidence" — the answer's accuracy survives, but the evidence behind it is quietly weakened.
Why This Is a "Proxy Objective Trap"
This finding is a clean instance of the "proxy objective trap."
Standard PTQ (Post-Training Quantization) optimizes reconstruction error between the quantized and full-precision models — making weights, activations, and KV cache as close as possible. The proxy quantity is token-level KL divergence — the difference between the two models' next-token distributions.
This proxy is reasonable in most settings: the closer the token distributions, the closer the behavior. But the paper exposes its blind spot: KL divergence treats all tokens equally, yet different tokens contribute vastly differently to "rationale credibility."
- If the word "therefore" is wrong, the rationale's credibility is unaffected.
- If "ST-segment elevation" is wrong, the entire rationale collapses.
- OSTQuant baseline (standard PTQ): preserves task accuracy, but answer-rationale support drops significantly. Answers remain correct, but rationales no longer cite the correct clinical evidence the way the full-precision model does.
- Explanation-aware method: preserves task accuracy while better retaining the full-precision model's answer choices and rationale-answer consistency.
- Still answer correctly (accuracy unchanged)
- Still produce fluent rationales (fluency unchanged)
- Yet cite different evidence — from "ST-segment elevation" to "age over 50"
- RAG retrieval relevance ≠ trustworthiness: a retrieved document being relevant doesn't make it reliable.
- V-JEPA pixel reconstruction ≠ rotation awareness: a model that reconstructs pixels doesn't necessarily understand rotation.
- Aggregated moral labels ≠ ground truth: aggregating multiple moral labels can deviate from genuine moral judgment.
- PTQ accuracy ≠ rationale support: unchanged answer accuracy doesn't mean rationales still support answers.
- The Faithfulness Cache depends on full-precision model quality. If the teacher's rationales are themselves flawed, the "behavior to preserve" is flawed too. Restricting to correctly-answered samples mitigates but doesn't eliminate this — the full-precision model can be right with a bad rationale.
- Evidence-token identification is not fully automatic. The paper uses KL-based clustering with hyperparameters and heuristics; different identification methods may yield different protection effects.
- The method is bound to OSTQuant. The explanation-aware objective is instantiated on OSTQuant; applicability to other PTQ methods (GPTQ, AWQ, SmoothQuant, etc.) is untested, though plausible in principle.
- Only medical QA was tested. Whether other explanation-critical domains (legal reasoning, scientific reasoning, code explanation) show the same problem is unaddressed, though the idea should generalize to any setting where rationale-answer consistency matters.
Standard PTQ treats these two token types the same. The majority generic discourse tokens dominate the objective, and the sparse but critical clinical evidence tokens are drowned out. The quantized model looks very similar to the full-precision model at the token level (small KL divergence), but diverges on the dimension that actually matters: whether the rationale supports the answer.
The proxy (token-level KL divergence) aligns with the true objective (rationale-answer consistency) on the training distribution, but fails under the distribution shift that quantization introduces — it faithfully protects what shouldn't be prioritized while ignoring what truly matters.
Method: Faithfulness Cache and Explanation-Aware Losses
The paper's solution has two steps.
Step 1: Build the Faithfulness Cache
Generate rationales for calibration samples from the full-precision (teacher) model, then for each sample:
1. Identify the evidence tokens supporting the answer — not all tokens, but the key tokens "without which the answer might change." 2. Record these evidence tokens alongside the corresponding answer behavior, forming a "Faithfulness Cache."
The cache includes only samples the full-precision model answered correctly — avoiding enshrining the model's own mistakes as "behavior to preserve."
Step 2: Explanation-Aware PTQ Objective
Building on OSTQuant (a transformation-based PTQ method), two additional loss terms are added:
1. Evidence-token protection loss: for evidence tokens in the Faithfulness Cache, increase their weight in the KL divergence — forcing the quantized model's distributions on these tokens closer to the full-precision model. 2. Rationale-conditioned answer behavior loss: protect not just the evidence tokens themselves, but "the model's predicted answer distribution given the evidence tokens" — ensuring the evidence-to-answer mapping isn't broken by quantization.
Together, these losses concentrate optimization pressure on the tokens that truly matter, rather than letting it be diluted by generic discourse words.
Experiments: Medical QA Under W4A4KV4 Quantization
The paper evaluates four 7B-8B medical and instruction-tuned models under W4A4KV4 (4-bit weights, 4-bit activations, 4-bit KV cache) on three medical QA benchmarks: MedExQA, MedExpQA, and ChallengeClinicalQA.
Key findings:
The structure of this result is notable: the two methods may be comparable on "answer accuracy," but differ clearly on "whether rationales support answers." If your evaluation only looks at accuracy, quantization seems fine; only when you evaluate rationale-answer consistency does the problem surface.
Why This Matters: The Evaluation Blind Spot in Medical AI
This finding has direct implications for medical AI evaluation.
Most current medical LLM evaluations use answer accuracy — the model picks the right option and it counts. This implicitly assumes: if the answer is right, the reasoning can be trusted.
But the paper exposes how fragile that assumption is. A quantized model may:
If doctors use rationales to understand the model's decision basis, the "clinical reasoning" they learn may be wrong. In medical settings this is no small matter — a doctor's understanding of the model's reasoning directly affects whether they trust it and how they act on its suggestions.
The paper's recommendation: in explanation-critical settings, PTQ evaluation should include rationale-answer consistency metrics, not just answer accuracy. This is a widely neglected evaluation dimension.
Broader Implications: The Generality of Proxy Failure
This finding echoes a broader pattern. In AI evaluation, "proxy failure" is a recurring theme:
All these traps share the same structure: we use a measurable proxy to evaluate a hard-to-measure true objective; the proxy aligns with the true objective on the training distribution, so we assume it detects the true objective. But when the distribution changes (quantization is a distribution shift), the proxy may faithfully protect what it actually measures — which may not be what you thought mattered.
The paper's contribution is not just a better PTQ method, but diagnosing a neglected evaluation blind spot — in explanation-critical settings, "is the answer right" and "is the reasoning right" are two different things, and standard PTQ objectives and evaluation metrics can both blind you to degradation in the latter.
Honest Assessment: Limitations
The method has notable boundaries:
A Deeper Question: What Counts as "Evidence"?
The method implicitly assumes some tokens in a rationale are "evidence" and others are "discourse." The distinction is based on KL clustering — grouping tokens by their contribution to quantization loss and identifying the class (clinical cue words) whose KL contribution share falls below their rationale share.
But "what is evidence" may be more complex than the paper assumes. In some contexts, "therefore" may genuinely be part of the evidence chain, marking a reasoning pivot; a seemingly generic discourse word may carry a key logical relation.
The clustering approach is a practical approximation, not a fundamental definition of evidence. A more principled approach might define evidence causally — which tokens, if substituted, would significantly change the answer's probability distribution? That approaches counterfactual reasoning, at higher computational cost. How to more principledly identify evidence tokens is a worthwhile question for follow-up work; the paper offers a practical answer, not the final one.
Conclusion
"When Quantization Preserves Accuracy but Not Evidence" is a diagnostic paper. It diagnoses a neglected problem (quantization preserving accuracy doesn't guarantee preserving evidence), offers a practical solution (explanation-aware PTQ), and reveals a broader evaluation blind spot (proxies can fail under distribution shift).
For model deployers: in explanation-critical settings, evaluating quantized models requires rationale-answer consistency, not just accuracy. Standard PTQ may degrade dimensions you can't see.
For AI evaluation researchers: this paper is another clean instance of the "proxy objective trap." When a proxy (token-level KL divergence) aligns with the true objective (rationale-answer consistency) on the training distribution, we assume it optimizes the true objective. But a proxy measures what it measures, not what you think it measures — and under distribution shift, the truth comes out.
For medical AI practitioners: rationale-answer consistency should become part of standard evaluation. Doctors need not just the answer but the why. If quantization quietly changes the "why" even when the "what" is unchanged, the model's clinical trustworthiness needs re-evaluation.
---
Paper: When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs arXiv: https://arxiv.org/abs/2609.24799 Code: https://github.com/dut0817/EAQuant Affiliation: University of Alberta