Imagine seeing this post on social media: "The Earth is flat, and NASA has been lying to us all along."
Is this a conspiracy theory? If the poster means it sincerely, yes. But what if they are mocking conspiracy theorists? What if it was posted in a satire group, alongside a flat-earth meme?
The same surface content can express endorsement, reasonable concern, criticism, sarcasm, or mockery. You cannot tell from the text alone — and this is the core challenge of conspiracy detection.
Klein, Shapira, and Hirsch at the Hebrew University, in *Agentic Detection of Online Conspiracies*, propose a solution: don't just analyze the text — infer the speaker's intent.
The Return of Speech Act Theory
The paper's theoretical foundation is Austin's 1962 speech act theory: a sentence doesn't just describe facts, it *does* things — endorsing, questioning, mocking, persuading. The same proposition, "vaccines are harmful," can be:
- Endorsement: the speaker genuinely believes vaccines are harmful
- Reasonable concern: evidence-based doubts about vaccine safety
- Criticism: opposition to anti-vaccine narratives
- Sarcasm: using irony to mock anti-vaxxers
- Mockery: imitating conspiracy rhetoric to make fun of it
The Agent Framework: Tool-Augmented Intent Inference
The authors built a LangGraph-based agent framework whose core idea is: don't just read the text — query context on demand. The agent can invoke three types of tools:
User history tool: queries the poster's past behavior. What have they posted before? Do they often post ironically? What topics do they follow?
Social network tool: queries the poster's social graph. Who do they follow? Who follows them? What communities are they in?
Post context tool: queries the conversation thread. Who are they replying to? What started the discussion?
The workflow: read the post, decide whether more context is needed; if so, pick the right tool; update the judgment based on results; possibly query again; finally classify.
This mirrors human judgment: when unsure what a sentence means, we check the person's profile, see who they're replying to, and look at their audience.
Key Findings: The Value of Context
The experiments reveal several interesting phenomena:
Overall gains are substantial. A text-only LLM achieves F1 of 0.535; the full agent framework reaches 0.73 — a 37% improvement. But the details are even more interesting:
User history is the most valuable single tool. Using user history alone (F1=0.698) beats feeding all available context to the model at once (F1=0.670). This shows that on-demand agent querying beats information bombardment — the model doesn't need all the information, it needs the right information at the right moment.
Context can sometimes hurt. 78% of the agent's errors were false positives — flagging non-conspiratorial content as conspiracies. The authors note: "sometimes text should be taken at face value; incomplete or over-reliant contextualization can distort interpretation."
Two-agent debate performed worst. Having two agents debate before judging (F1=0.558) barely beat text-only. The debate framework introduced more noise than insight.
The Challenge of an "Adversarial" Dataset
The authors stress that their dataset is "adversarial" — the non-conspiracy samples contain many posts whose surface language resembles conspiracy claims (sarcasm, mockery, criticism). This is closer to real-world conditions.
On easy datasets (clear separation between conspiracy and normal posts), text-only classifiers do fine. But in the real world, the hard part isn't recognizing an obvious claim like "the Earth is flat" — it's distinguishing "the Earth is flat (sincere)" from "the Earth is flat (sarcastic)."
On the adversarial dataset, the full agent reduced false negatives from 58 to 30 and false positives from 142 to 76. Context helped fix both error types: confirming genuine conspiracies and ruling out superficially similar non-conspiracies.
Token Economics: Agents vs. Context Stuffing
The authors ran an interesting "token economy" analysis: compared to dumping all context into the model at once, how cost-effective is the agent framework?
Result: the agent framework achieved better performance with fewer input tokens. On-demand querying avoids irrelevant information and saves compute. This aligns with the "judgment-gate decoupling" idea — deciding what information you need before fetching it works better than receiving everything upfront.
My Take
The paper's core insight: content classification requires intent inference, intent inference requires context, and context acquisition requires an agent. This is not simple RAG — RAG "retrieves relevant information to append to the prompt," while an agent "actively decides what information is needed, where to find it, and how to use it."
The finding that 78% of errors are false positives deserves reflection. Faced with ambiguous content, the model leans toward "presumption of guilt" — better to over-flag than miss. This matches human moderator bias: when told to "find conspiracies," you become suspicious of everything. But the correct default should be "presumption of innocence" — take text at face value unless context proves otherwise.
The worst-performing two-agent debate setup is also worth remembering. Two agents debating is not "two experts discussing" — it's more like "two stubborn people arguing," reinforcing their positions rather than converging on the truth. As with cross-domain analogies to Milgram's obedience experiments, multi-agent systems are not automatically better than single agents; the interaction structure design is what matters.
---
Paper: Klein, E., Shapira, B., & Hirsch, L. (2026). Agentic Detection of Online Conspiracies. arXiv:2609.30250 Link: https://arxiv.org/abs/2609.30250