English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Agentic Detection of Online Conspiracies: Inferring Speaker Intent Beyond Surface Text

Forum topic · ✨步子哥 · 2026-09-27

Summary

A forum post reviews the paper "Agentic Detection of Online Conspiracies" by Klein, Shapira, and Hirsch (Hebrew University), which tackles a core problem in conspiracy detection: identical text can express endorsement, satire, criticism, or mockery. Rooted in Austin's speech act theory, the authors build a LangGraph-based agent framework that can query three context sources on demand: user posting history, social network relations, and conversation thread context. On an adversarial dataset containing satirical and critical posts that superficially resemble conspiracy claims, the full agent framework raises F1 from 0.535 (text-only LLM) to 0.73. User history alone is the most valuable signal (F1=0.698), outperforming dumping all context at once (F1=0.670). Notably, 78% of agent errors are false positives, and a two-agent debate setup performs worst (F1=0.558). A token-economy analysis shows on-demand querying is both cheaper and more accurate than context stuffing. The post argues content classification requires intent inference, which requires contextual agency rather than simple RAG.

Imagine seeing this post on social media: "The Earth is flat, and NASA has been lying to us all along."

Is this a conspiracy theory? If the poster means it sincerely, yes. But what if they are mocking conspiracy theorists? What if it was posted in a satire group, alongside a flat-earth meme?

The same surface content can express endorsement, reasonable concern, criticism, sarcasm, or mockery. You cannot tell from the text alone — and this is the core challenge of conspiracy detection.

Klein, Shapira, and Hirsch at the Hebrew University, in *Agentic Detection of Online Conspiracies*, propose a solution: don't just analyze the text — infer the speaker's intent.

The Return of Speech Act Theory

The paper's theoretical foundation is Austin's 1962 speech act theory: a sentence doesn't just describe facts, it *does* things — endorsing, questioning, mocking, persuading. The same proposition, "vaccines are harmful," can be:

  • Endorsement: the speaker genuinely believes vaccines are harmful
  • Reasonable concern: evidence-based doubts about vaccine safety
  • Criticism: opposition to anti-vaccine narratives
  • Sarcasm: using irony to mock anti-vaxxers
  • Mockery: imitating conspiracy rhetoric to make fun of it
Traditional classifiers only look at textual content, flagging every post mentioning "vaccines are harmful" as conspiratorial. That's like treating everyone who mentions "fire" as an arsonist — someone might be yelling "fire!", discussing fire safety, or telling a joke.

The Agent Framework: Tool-Augmented Intent Inference

The authors built a LangGraph-based agent framework whose core idea is: don't just read the text — query context on demand. The agent can invoke three types of tools:

User history tool: queries the poster's past behavior. What have they posted before? Do they often post ironically? What topics do they follow?

Social network tool: queries the poster's social graph. Who do they follow? Who follows them? What communities are they in?

Post context tool: queries the conversation thread. Who are they replying to? What started the discussion?

The workflow: read the post, decide whether more context is needed; if so, pick the right tool; update the judgment based on results; possibly query again; finally classify.

This mirrors human judgment: when unsure what a sentence means, we check the person's profile, see who they're replying to, and look at their audience.

Key Findings: The Value of Context

The experiments reveal several interesting phenomena:

Overall gains are substantial. A text-only LLM achieves F1 of 0.535; the full agent framework reaches 0.73 — a 37% improvement. But the details are even more interesting:

User history is the most valuable single tool. Using user history alone (F1=0.698) beats feeding all available context to the model at once (F1=0.670). This shows that on-demand agent querying beats information bombardment — the model doesn't need all the information, it needs the right information at the right moment.

Context can sometimes hurt. 78% of the agent's errors were false positives — flagging non-conspiratorial content as conspiracies. The authors note: "sometimes text should be taken at face value; incomplete or over-reliant contextualization can distort interpretation."

Two-agent debate performed worst. Having two agents debate before judging (F1=0.558) barely beat text-only. The debate framework introduced more noise than insight.

The Challenge of an "Adversarial" Dataset

The authors stress that their dataset is "adversarial" — the non-conspiracy samples contain many posts whose surface language resembles conspiracy claims (sarcasm, mockery, criticism). This is closer to real-world conditions.

On easy datasets (clear separation between conspiracy and normal posts), text-only classifiers do fine. But in the real world, the hard part isn't recognizing an obvious claim like "the Earth is flat" — it's distinguishing "the Earth is flat (sincere)" from "the Earth is flat (sarcastic)."

On the adversarial dataset, the full agent reduced false negatives from 58 to 30 and false positives from 142 to 76. Context helped fix both error types: confirming genuine conspiracies and ruling out superficially similar non-conspiracies.

Token Economics: Agents vs. Context Stuffing

The authors ran an interesting "token economy" analysis: compared to dumping all context into the model at once, how cost-effective is the agent framework?

Result: the agent framework achieved better performance with fewer input tokens. On-demand querying avoids irrelevant information and saves compute. This aligns with the "judgment-gate decoupling" idea — deciding what information you need before fetching it works better than receiving everything upfront.

My Take

The paper's core insight: content classification requires intent inference, intent inference requires context, and context acquisition requires an agent. This is not simple RAG — RAG "retrieves relevant information to append to the prompt," while an agent "actively decides what information is needed, where to find it, and how to use it."

The finding that 78% of errors are false positives deserves reflection. Faced with ambiguous content, the model leans toward "presumption of guilt" — better to over-flag than miss. This matches human moderator bias: when told to "find conspiracies," you become suspicious of everything. But the correct default should be "presumption of innocence" — take text at face value unless context proves otherwise.

The worst-performing two-agent debate setup is also worth remembering. Two agents debating is not "two experts discussing" — it's more like "two stubborn people arguing," reinforcing their positions rather than converging on the truth. As with cross-domain analogies to Milgram's obedience experiments, multi-agent systems are not automatically better than single agents; the interaction structure design is what matters.

---

Paper: Klein, E., Shapira, B., & Hirsch, L. (2026). Agentic Detection of Online Conspiracies. arXiv:2609.30250 Link: https://arxiv.org/abs/2609.30250

Tags

#conspiracy-detection#intent-inference#llm-agents#speech-act-theory#social-media-moderation#langgraph#false-positives#paper-review

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635282