English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Evidence-Type Competition: LLMs Learn Causal Effect Magnitudes from Interventions but Copy Direction from Observations

Forum topic · ✨步子哥 · 2026-08-04

Summary

A 2026 Tsinghua University study (arXiv:2607.29484) tested the intuitive hypothesis that increasing the proportion of interventional data in pretraining improves LLM causal reasoning. In synthetic Simpson's paradox worlds where observational correlations oppose true causal effects, the hypothesis was refuted: as intervention data ratio alpha rose from 0 to 1.0, direction accuracy stayed flat (0.2-0.5), yet predicted effect magnitudes grew monotonically (0.169 to 0.384). This magnitude-direction duality shows models learn effect size from interventional data while copying effect direction from observational context. Inference-time interventions reveal the causal capability exists in weights but is suppressed by contextual observational evidence: erasing observational records corrects direction (ratio_true from -0.09 to +0.56), and activation patching localizes the switch to observation-evidence rows in middle layers. The paper also identifies a 26% sampling noise floor in probe-based causal evaluation, reducible to 9% via evidence averaging, and a positive-effect prior learned from CLadder. Key takeaway: training data mix is not the controlling knob; evidence-type competition at inference time is.

Paper: https://arxiv.org/abs/2607.29484

Imagine teaching a child about causality with two kinds of data:

  • Observational data: in the natural state, every time ice cream sales rise, heatstroke rates rise too — but both are driven by summer, not by ice cream.
  • Interventional data: randomly assigning people to eat ice cream in winter shows no change in heatstroke — this reveals the true causal effect.
  • Conventional wisdom says interventional (randomized) data is the gold standard, so feeding a model more of it should improve causal reasoning. A July 2026 Tsinghua University team found otherwise, naming the phenomenon magnitude–direction duality.

    H1: An Intuition Never Rigorously Tested

    The paper tests hypothesis H1: increasing the proportion of interventional data in pretraining improves causal reasoning. This assumption underpins Pearl's ladder of causation and causal inference methodology generally, but the authors found it had never been tested under strictly controlled conditions:

  • Known causal graph structure
  • Computable causal ground truth
  • Leak-free train/test protocols
  • Controllable confounding strength
  • Since real-world data cannot satisfy these, the authors built a synthetic world generator with known causal structure.

    Simpson's Paradox Worlds: Observation vs. Intervention in Opposition

    The experiments use worlds where:

  • Observational correlation is negative: Corr_obs(X,Y) < 0
  • True causal effect is positive: eff(X→Y) > 0
  • This creates a causal trap: a model relying on observational data infers exactly the wrong direction. Six experiments sweep the interventional data ratio alpha from 0 (pure observational) to 1.0 (pure interventional).

    H1 Refuted: Magnitude Learned, Direction Not

  • Direction accuracy across six alpha levels: 0.5 / 0.3 / 0.2 / 0.4 / 0.4 / 0.4 — no monotonic improvement. H1 is refuted.
  • Simpson slope magnitude grows monotonically with alpha: 0.169 → 0.384.
  • Direction stays wrong — always negative, tracking observational data.
  • > The model learns the magnitude of causal effects from interventional data, but copies the direction from observational context.

    Key Finding: The Switch Is in Context, Not in Weights

    Experiment 1: Inference-time intervention

  • Doubling interventional evidence in context → no effect; direction stays wrong.
  • Erasing observational evidence from context → direction immediately corrects! ratio_true jumps from -0.09 to +0.56, matching purely interventional-trained models.
  • The causal reasoning capability is in the weights; observational evidence in context switches it off.

    Experiment 2: Content manipulation

  • 1 observational record → almost no suppression
  • 2 records → dilution begins
  • 4 records → full capture (direction flips)
  • Replacing record content while preserving format releases suppression — the switch is content-mediated, not format-mediated.

    Experiment 3: Activation patching

    Activation patching localizes the switch to observation-evidence rows in middle layers. Erasing them restores causal reasoning.

    Sampling Noise Floor: 26% of "Correct" Answers Are Luck

    A methodological contribution: when evaluating with probes (interventions on query variable X, observing responses), even a perfect least-squares interpolator commits 26% direction errors, because regressions over 4 sampled intervention points are inherently noisy. Evidence averaging — sampling multiple evidence sets per query — reduces the noise floor from 26% to 9%. Warning for all probe-based causal evaluations: your evaluation may be noisier than your model.

    CLadder Audit: A Learned Positive-Effect Prior

    An external audit on CLadder finds a positive-effect prior — a tendency to predict positive causal effects — with two layers:

  • Surface: removable via sign-randomized retraining (in-distribution)
  • Deep: a positive-effect default persists out-of-distribution
  • Deeper Implications

    1. "Capability lives in the weights; the switch lives in the context." This is not missing capability but evidence-selection failure: two statistical evidence types (observational vs. interventional) compete, and the model defaults to the wrong one. 2. Training data ratio is not a universal knob. When the root cause is inference-time evidence competition, adjusting the training mix cannot help — you must intervene on context at inference time. 3. Another evaluation blind-spot law. Looking only at magnitude or aggregate accuracy misses the direction-wrong failure mode; the 26% noise floor means uncontrolled causal evaluations may be pure noise. 4. Simpson's paradox as a diagnostic tool. Worlds where observation and intervention disagree expose the model's evidence-selection strategy — in direction-consistent worlds, you cannot tell whose "homework" the model is copying.

    Limitations

    1. Synthetic worlds: a 25.7M-parameter model on synthetic causal structures; real LLMs on real data may differ. 2. 0.93B validation: rate-level persistence holds at 0.93B parameters, but absolute interventional gains shrink 4x — does the phenomenon survive in larger models? 3. Single causal structure: only linear–nonlinear mixed mechanisms tested; feedback loops and temporal dependencies untested.

    Conclusion

    "More good data" does not equal "better learned capability." When two evidence types compete, data ratio is not the referee — context is. The model is not a black box fixed by training data, but a system continuously performing evidence selection at inference time.

    ---

    Paper: https://arxiv.org/abs/2607.29484

    FAQ

    Q1: Who is this for? Practitioners, researchers, and students interested in AI, machine learning, and deep learning.

    Q2: What are the core points?

  • H1: an intuition never rigorously tested
  • Simpson's paradox worlds where observation and intervention disagree
  • H1 refuted: magnitude learned, direction copied from observation
Q3: Is there open-source code? See the links in the article body.

Tags

#causal-inference#large-language-models#evidence-type-competition#simpsons-paradox#activation-patching#model-evaluation#training-data#interpretability

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503933