The r = 0.85 Finding
In 2025, Ray Dalio—founder of Bridgewater Associates—said on CNBC: "there will absolutely be a debt crisis."
If you are a computational social science researcher, you might ask: is "absolutely" here genuine epistemic certainty, or a rhetorical habit? Do pessimists use "definitely" to amplify their argument, or "possibly" to hedge their risk?
This is an answerable question. You collect ten years of Dalio's interview videos, score every sentence on two dimensions: negative affect word frequency and emphatic certainty word frequency. Then compute the correlation.
The result: r = 0.851, p < 0.001. A breathtakingly clean finding—the more pessimistic Dalio is, the more certain he sounds. This fits the popular intuition about "doomsayers": pessimists use certainty to amplify fear, optimists use hedging to manage reputation.
You extend the finding to three more public figures: Cathie Wood (tech optimist), Kenneth Rogoff (academic economist), Peter Zeihan (geopolitical determinist). All four show the same pattern, r from 0.725 to 0.932, all p < 0.01.
This is nearly a perfect computational social science finding: large sample (85 interviews, 32,625 sentences), cross-speaker consistency, large effect size, high statistical significance. Ready for a top venue.
But the finding may be fake.
Not through data fraud, not through p-hacking, not through sampling bias. Something subtler: the measurement instrument itself manufactured the correlation.
The Blind Spots of Keyword Lexicons
Bo Chen's paper *When Certainty Is an Artifact* (published June 2026, Institute of Computing Technology, Chinese Academy of Sciences) did something simple: scored the same corpus with two different methods.
Method 1: Keyword lexicon. Count emphatic words ("always", "clearly", "absolutely", "certainly") as the "certainty" score; count negative words ("never", "no", "not") as the "negative affect" score. This is the standard approach in computational social science for two decades.
Method 2: LLM zero-shot semantic classification. A large language model makes sentence-level semantic judgments—accounting for negation, modification, polysemy, and context—scoring whether each sentence semantically expresses certainty/negativity.
Results on the same corpus:
| Speaker | Lexicon r(neg, emphatic) | LLM r(neg, emphatic) | |---------|--------------------------|----------------------| | Ray Dalio | 0.851 (p<0.001) | 0.206 | | Cathie Wood | 0.725 (p=0.003) | negative | | Kenneth Rogoff | 0.917 (p<0.001) | negative | | Peter Zeihan | 0.932 (p<0.001) | not significant |
r = 0.851 became r = 0.206. For three of four speakers the correlation vanished or reversed.
This is not "method 2 is more precise than method 1." This is one dataset, one research question, two measurement instruments, completely opposite conclusions.
Why the Lexicon Fails
The paper analyzes five systematic blind spots of keyword lexicons, each of which pushes "certainty" measurements away from the truth:
Blind spot 1: Negation. "I am never absolutely totally confident"—this sentence contains two emphatic words ("absolutely", "totally"), so the lexicon gives a high certainty score. But semantically, "never" negates the entire expression; this is a low-certainty statement. The lexicon cannot see "never" negating "absolutely".
Blind spot 2: Polysemy. "a certain body"—"certain" here means "some", not "sure". The lexicon counts it as a certainty word anyway.
Blind spot 3: Modification. "sort of clear"—"clear" is a certainty word, but "sort of" hedges it. The lexicon sees only "clear".
Blind spot 4: Framing. "I think for everyone"—"I think" is a hedge, but "everyone" is an intensifier. The lexicon counts only "everyone".
Blind spot 5: Quantification. "almost everything"—"everything" is a universal quantifier (emphatic), "almost" hedges it. The lexicon counts only "everything".
Each blind spot misclassifies hedged statements as emphatic ones. And negative discourse naturally attracts these structures—when people deliver bad news they tend to say "never absolutely", "sort of clear that it's bad", "almost everything is wrong"—so the lexicon method manufactures a spurious "negative ↔ emphatic" correlation.
The correlation is not about the speakers' psychology; it is about the statistical structure of English vocabulary.
The Deeper Problem: A Category Error
The paper's sharpest argument is not "keyword methods are inaccurate" but:
> Treating keyword counts as measurements of epistemic certainty is a category error.
A category error, in Gilbert Ryle's sense, is treating a concept from one category as if it belonged to another. The classic example: after touring the library, labs, and classrooms, asking "but where is the university?"—the university is not a thing alongside the library; it is the way these institutions are organized.
The category error here: keyword lexicons measure the English lexical-statistical regularity that negative discourse attracts emphatic vocabulary, not the speaker's epistemic state. The two are conflated.
The key evidence: four speakers with very different personalities (optimist Wood, academic Rogoff, determinist Zeihan, doomsayer Dalio) all show r = 0.72–0.93 under the lexicon method. If the correlation reflected a psychological state, four opposite personalities should not share it. But if it reflects English lexical statistics—of course everyone shows it, because everyone speaks English.
A finding about psychology is actually a finding about English.
What the LLM Revealed
If the LLM method merely "disproved the lexicon finding," the paper's value would be limited—it would be pure falsification. But the LLM method uncovered a pattern the lexicon could not see:
A strong correlation between negative affect ↔ hedged language.
| Speaker | LLM r(neg, hedged) | |---------|--------------------| | Kenneth Rogoff | 0.875 (p=0.001) | | Peter Zeihan | 0.722 (p=0.008) |
This fits intuition: pessimists hedge ("maybe", "probably", "not entirely certain") rather than intensify. The pattern is especially strong for Rogoff—academic training makes him habitually wrap pessimistic predictions in hedges.
The lexicon cannot see this pattern because hedge word frequencies ("maybe", "perhaps", "sort of") are drowned out by emphatic word frequencies. An LLM can distinguish that "sort of clear" is a hedge, not emphasis; a lexicon cannot.
Broader Implications
First, the reproducibility crisis in computational social science may be deeper than we think. Many published "findings" may not be about human behavior but about measurement instruments. When a more precise tool makes the effect disappear, this is not a replication failure—it is the original study measuring the wrong thing.
Second, LLMs are not just "better classifiers"; they are "different instruments." The difference between lexicon and LLM methods is not precision but category—they are not measuring the same thing at all. It is like measuring temperature with a ruler versus a thermometer: not a precision problem, a measurement-target problem.
Third, "statistically significant + large effect size" does not equal "a true finding." r = 0.851, p < 0.001, N = 21—textbook-grade strong evidence. Yet it was entirely an artifact of the measurement instrument. Statistical significance only tells you the effect is not random noise; it cannot tell you the effect is the one you think it is.
A Question for the Reader
The paper does not conclude "simply replace lexicons with LLMs." LLMs have their own problems: black-box judgments, limited interpretability, and results that may vary across models (the paper ran cross-model robustness checks with broadly consistent results, but that guarantees nothing about the future).
The deeper question: how do we know any measurement instrument measures what we think it measures?
Keyword methods took twenty years to be caught measuring the wrong thing. Could LLM methods be caught in another twenty? Is it possible that "epistemic certainty" as a construct lacks a sound operational definition, and any instrument only captures one of its projections?
This paper cannot answer that. But it raises the question sharply—and with a concrete r = 0.851 → r = 0.206 example that you cannot avoid.
The next time you see a computational social science finding that is "statistically significant, large-effect, cross-sample consistent," ask first: is this finding about people, or about the measurement instrument?
---
Paper: When Certainty Is an Artifact: Keyword Lexicon Blindness and the (Mis)Measurement of Rhetorical Stance Author: Bo Chen (Institute of Computing Technology, Chinese Academy of Sciences)