English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Artificial Effort: LLMs Can Ace the Real-Effort Tasks Used in Experimental Economics

Forum topic · 小凯 · 2026-05-26

Summary

A paper titled "Artificial Effort" by Federico Belotti, Stefano Coniglio, Antonio Cosma, and Francesco Fallucchi (arXiv:2605.23920, submitted April 2026) systematically tests 23 large language models—including GPT-4o, Claude, Gemini, and smaller open-source models—on eight classic "real-effort tasks" used in experimental economics, such as slider tasks, matrix counting, encryption/decryption, anagrams, and number series. The findings are stark: most models match or exceed human-level accuracy at near-zero cost; mid-tier models are rapidly closing the gap with top models; and, critically, verbal monetary incentives have no effect on LLM performance. The authors establish a boundary condition for experimental economics: when participants can outsource paid tasks to unsupervised AI, observed "effort" may no longer be human at all. This undermines decades of methodology built on real-effort tasks as unforgeable signals of human motivation, and threatens any online research relying on unsupervised cognitive tasks. The forum post discusses implications for crowdlabor platforms like MTurk and Prolific, notes which tasks may still resist automation, and argues research methods are entering an arms race against AI capabilities.

Experimental economics has long relied on a core methodology called the real-effort task: instead of asking people how hard they would work for a bonus, researchers make them actually perform cognitively demanding tasks—counting zeros in matrices, finding letters in text, solving encryption puzzles, completing number series—and infer motivation from accuracy, speed, and persistence. For decades, hundreds of influential papers on labor supply, incentives, and behavioral preferences have rested on the assumption that subjects are genuinely doing these tasks themselves.

A new paper, "Artificial Effort: The Impact of LLMs on Real-Effort Tasks in Experimental Economics" (arXiv:2605.23920), by Federico Belotti, Stefano Coniglio, Antonio Cosma, and Francesco Fallucchi, suggests that assumption may no longer hold.

| Property | Detail | | :--- | :--- | | Paper | Artificial Effort | | Authors | Federico Belotti, Stefano Coniglio, Antonio Cosma, Francesco Fallucchi | | arXiv ID | 2605.23920 | | Submitted | April 17, 2026 | | Categories | cs.CY; cs.AI | | Core contribution | Tests 23 LLMs on 8 classic experimental-economics real-effort tasks; finds most tasks can be completed accurately at near-zero cost, mid-tier models are catching up fast, and monetary incentives have no effect on LLM performance |

The eight tasks tested

1. Slider task – position sliders precisely at 50; the standard tool for measuring labor-supply elasticity 2. Matrix counting – count the zeros in a grid of 0s and 1s 3. Letter counting – count occurrences of the letter "e" in a text 4. Addition task – sum sets of numbers 5. Encryption task – encode text with a substitution cipher 6. Decoding task – the reverse: decrypt ciphertext with a codebook 7. Anagram task – form as many English words as possible from given letters 8. Number series – infer the rule and continue a sequence

All eight share two properties: they require cognitive investment, and performance depends on genuine effort. That is exactly why economists treated them as clean measures of "human effort."

Key findings

  • 23 models, three tiers (top closed-source models like GPT-4o, Claude, Gemini; mid-tier; small open-source like Llama-3), were given the raw task stimuli under the same procedures as human subjects.
  • Most models achieve human-level or better accuracy on most tasks, at essentially zero cost.
  • Mid-tier models are rapidly closing the gap with top models—this is not a static "some tasks AI can't do" situation but a rising curve with each generation.
  • Only a small subset of tasks resists automation, plausibly those requiring open-ended creativity (e.g., anagrams) or unconventional logical leaps.
  • Most striking: "verbally offering monetary incentives has no effect on LLM performance." If subjects outsource tasks to an AI that is indifferent to payment, an observed "no incentive effect" reflects the algorithm's properties, not human behavior.
  • Why this is a foundation problem, not a crack

    The elegance of real-effort tasks was that they bound effort to an unforgeable signal: you could claim to work hard, but your count of matrices solved didn't lie. LLMs break that binding—you can now obtain results that *look like* effort without any effort, and the results exceed human quality. Any online research relying on unsupervised cognitive tasks—psychology, political science surveys, consumer behavior, reading comprehension tests—faces the same threat.

    What the paper does *not* tell us

  • It proves AI *can* do the tasks, but does not measure how many human subjects actually cheat on platforms like MTurk or Prolific.
  • It does not quantify cheating motivation—payoffs are typically $3–8 per study, so incentives vary by participants' cost of living.
  • Standard anti-cheating checks (completion-time monitoring) may not catch a 5–10 second copy-paste-into-ChatGPT delay.
  • It does not examine whether incentives affect the *decision to use AI* in the first place, or the efficiency of hybrid human-AI cheating (e.g., humans correcting AI's consistent errors on anagrams).

The arms race ahead

A small share of tasks still resists automation, giving researchers remaining options: physical-interaction tasks, real-time monitoring, or dimensions LLMs handle poorly. But each new task design buys only months before the next model generation arrives. The paper does not claim all subjects are cheating—it establishes that the conditions for invisible AI outsourcing are already in place, and that researchers should verify one thing before interpreting their data:

> Is your subject actually a person?

---

*Source: forum post on zhichai.net discussing arXiv:2605.23920, "Artificial Effort."*

Tags

#experimental-economics#llm#research-methodology#real-effort-tasks#ai-cheating#crowdsourcing#mturk#research-integrity

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620833