Imagine teaching a child about causality with two kinds of data:
- Observational data: naturally, ice cream sales and heatstroke rates rise together—but only because summer causes both.
- Interventional data: randomly assign people to eat ice cream in winter, and heatstroke doesn't change—this is the true causal effect.
- Known causal graph structure (who causes whom)
- Computable causal ground truth
- Leakage-free train/test protocols
- Controllable confounding strength
- Negative observational correlation: Corr_obs(X,Y) < 0
- Positive true causal effect: eff(X→Y) > 0
- Doubling interventional evidence (more interventional samples) → no effect, direction still wrong
- Erasing observational evidence (deleting observational data from context) → direction immediately corrected! ratio_true jumps from -0.09 to +0.56, matching models trained purely on interventions
- 1 observational record → almost no suppression
- 2 records → dilution begins
- 4 records → full capture (direction reversed)
- Surface layer: removable via sign-randomized retraining (in-distribution)
- Deep layer: an out-of-distribution default toward positive effects persists
Common sense says interventional data (randomized experiments) is the gold standard for causal inference. The more interventional data a model sees, the better its causal understanding should be.
But what if you found that as a model sees more interventional data, its estimated effect *magnitudes* get more accurate while its *direction* keeps following the observational data—insisting ice cream causes heatstroke, only downgraded from '100% of cases' to '50%'?
This is a real phenomenon, dubbed magnitude–direction duality, reported by a Tsinghua University team in a July 2026 arXiv paper.
Paper: https://arxiv.org/abs/2607.29484
The H1 Hypothesis: An Untested Intuition
The paper tests H1: increasing the proportion of interventional data in pretraining should improve causal reasoning.
The assumption seems obvious—Pearl's causal hierarchy places intervention above observation, and the field's methodology rests on 'randomized experiments beat observational studies.' In the LLM era, many extend this to training data: feed more interventional data, get better causal models.
Yet the authors found this hypothesis had never been tested under strictly controlled conditions, meaning:
These are nearly impossible with real-world data. So the paper builds a synthetic world generator—small worlds with known causal structures—and tests whether models can learn from them.
Simpson Paradox Worlds: Observation and Intervention Disagree
The experiment's clever design: a set of Simpson paradox worlds satisfying:
Directions are opposite. This is a causal trap: a model relying only on observational data gets the causal direction completely wrong.
Six experiments vary the interventional data fraction α from 0 (pure observation) to 1.0 (pure intervention). If H1 holds, higher α should yield better directional accuracy.
H1 Falsified: Magnitude Learned, Direction Not
The results are surprising.
Directional accuracy across the six α levels: 0.5 / 0.3 / 0.2 / 0.4 / 0.4 / 0.4. No monotonic trend. H1 is falsified.
But the story deepens:
The magnitude of the Simpson slope grows monotonically with α: 0.169 → 0.384. The model genuinely learns *that* a causal effect exists.
But the direction stays wrong—consistently negative, tracking the observational data. Hence 'magnitude–direction duality':
> The model learns the magnitude of the causal effect from interventional data, but the direction is copied from observational context.
It's like a physics student who knows 'gravity acts' (magnitude) but draws the arrow backward—because they keep copying the observational data's homework.
Key Finding: The Switch Is in Context, Not in Weights
If training-data ratio isn't the deciding factor, what is? The authors ran a series of inference-time interventions reminiscent of neuroscientists' inactivation experiments:
Experiment 1: Inference-time intervention
Conclusion: the causal reasoning capability was in the weights all along; contextual observational evidence was switching it off.
Experiment 2: Content manipulation
Moreover, replacing the records' content (keeping format, removing actual values) releases the suppression. The switch is content-mediated, not format-mediated—the model reads the *meaning* of observational data, not its appearance.
Experiment 3: Activation patching
Using activation patching, the authors localize the switch: observational evidence rows at mid layers are critical. Erasing them restores causal reasoning—consistent with the context-based switch, physically located in mid-layer attention.
The Sampling Noise Floor: 26% of 'Correct' Answers Are Luck
A methodological contribution worth highlighting: the sampling noise floor of probe-based evaluation.
If you probe a model's causal reasoning by intervening on a query variable X and observing responses, even a perfect least-squares interpolator makes 26% directional errors—because probes are randomly sampled. With only 4 intervention points sampled once each, the regression slope has inherent randomness: even with a positive underlying effect, the fitted slope can be negative.
It's like inferring whether a coin is fair from 4 flips—too few, too much variance. The fix is evidence averaging: sample multiple evidence sets per query. Averaging 8 sets drops the noise floor from 26% to 9%.
Warning for all probe-based causal evaluations: your evaluation may be noisier than your model.
CLadder Audit: A Learned 'Positive Effect Prior'
An external audit on CLadder (a causal reasoning benchmark) reveals a deeper issue: models learn a positive effect prior—a tendency to predict positive causal effects—with two layers:
Why This Matters
1. 'Capability lives in weights, the switch lives in context'
The paper's essence in one line. The model isn't incapable of causal reasoning—it was trained to have that capability—but contextual observational evidence suppresses it. This is not a capability deficit; it is an evidence-selection failure. Mirroring the 'knowledge in parameters, triggered by prompts' consensus, the causal-reasoning version is sharper: two types of statistical evidence (observational vs. interventional) compete, and the model defaults to the wrong one.
2. Training data ratios are not a universal knob
Many believe adjusting training-data proportions is a general fix for model capability. This paper says: not necessarily. When the root cause is inference-time evidence competition, changing training ratios won't help—you must intervene on context at inference time.
3. Another evaluation blind spot
If you only measure magnitude accuracy or overall accuracy, you miss the 'wrong direction' failure mode. The 26% noise floor means uncontrolled causal evaluations may be pure noise. Like QuantiBias and Progressive Cramming before it, this paper shows: you optimize what you measure; problems hide where you don't.
4. Simpson's paradox as a training-evaluation tool
Turning Simpson's paradox from a statistical trap into a causal diagnostic tool is a methodological innovation. When observational and interventional directions conflict, the model's evidence-selection strategy is exposed—in direction-consistent worlds, you can't tell whose homework it's copying.
Limitations
1. Synthetic worlds: 25.7M-parameter small models on synthetic causal structures; real LLMs on real data may behave differently. 2. 0.93B validation: rate-level persistence verified at 0.93B parameters, but absolute interventional benefit shrank 4×—does the phenomenon survive at larger scale? 3. Single causal structure: only linear–nonlinear mixed mechanisms tested; feedback loops and temporal dependencies remain untested.
Conclusion
The biggest lesson: 'more good data' does not equal 'better-learned capability.' When two evidence types compete, the data ratio isn't the referee—context is.
A model is not a black box determined by training data, but a system continuously selecting among evidence types at inference time. Understanding this is essential for building genuinely reliable AI systems.
---
Paper: https://arxiv.org/abs/2607.29484