English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Evaluating AI Beyond Benchmarks: Generative AI as a Sociotechnical System

Forum topic · 小凯 · 2026-05-04

Summary

A discussion of the paper "Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechnical Systems" (arXiv:2604.20545, 2026) by Rebecca L. Johnson, arguing that leaderboard-style scores measure far less than they appear to. The post critiques two dominant evaluation paradigms: functionalism, which treats models as isolated predictors judged by benchmark scores, and prescriptivism, which judges models against assumed universal human values. Both, it argues, treat AI as an object rather than a process. A sociotechnical lens reframes training data, annotation, deployment, and evaluation itself as products of social choices and power relations. The paper advocates pluralist evaluation: developers, policymakers, end users, and affected communities each hold distinct, complementary criteria—reasoning, safety, usability, dignity—that cannot substitute for one another. Echoing Feynman on observation, the post warns that benchmarks shape optimization: what we measure is what models pursue, potentially at the cost of creativity, common sense, or ethical judgment. It closes with reflective questions for anyone evaluating AI systems.

> Paper: Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechnical Systems > Author: Rebecca L. Johnson > arXiv: 2604.20545 | 2026-04-27

1. The World Where Leaderboards Rule

Open any AI evaluation site and you will see:

  • GPT-5.2: 92.3
  • Claude 4: 91.7
  • Gemini 3: 90.1
  • Llama 4: 88.5
  • These numbers look objective, scientific, unimpeachable. But the question is: what do they actually measure?

    The paper issues a fundamental challenge: evaluating generative AI is not just a technical problem—it is a sociotechnical one.

    2. Two Flawed Evaluation Paradigms

    The research critiques two mainstream approaches:

    1. Functionalism

  • Treats models as isolated predictors
  • Runs benchmarks; high scores mean "good"
  • Problem: ignores how models are trained, deployed, and used
  • 2. Prescriptivism

  • Evaluates what models "should" be
  • Uses human values as the standard
  • Problem: whose values? Which culture? Which era?
  • The shared blind spot of both paradigms: they treat models as "things" to be measured, rather than understanding AI systems as "processes."

    3. The Sociotechnical Systems Perspective

    What is a "sociotechnical system"?

    > AI is not isolated software. It is a complex system jointly constituted by technology, people, institutions, culture, and history.

    This means:

  • Training data is not "raw material" but the product of a social process (who decides what data to collect?)
  • Annotation is not "objective truth" but a collection of human judgments (who labeled it, and by what standards?)
  • Deployment is not a mere "application of technology" but an embodiment of power relations (who decides where AI is used?)
  • Evaluation is not "scientific measurement" but an expression of value choices (what do we choose to measure, and why?)
  • AI evaluation itself shapes AI development.

    4. Why Pluralism Is Necessary

    The paper argues for pluralist evaluation—not seeking one "correct" evaluation method, but acknowledging that:

  • Different stakeholders have different evaluation needs
  • Different cultural backgrounds carry different value standards
  • Different application contexts define success differently
  • For example:

  • Developers may care about "reasoning ability"
  • Policymakers may care about "safety and fairness"
  • End users may care about "usefulness and ease of use"
  • Affected communities may care about "autonomy and dignity"
These evaluations do not replace one another. They are complementary—and all necessary.

5. A Feynman-Style Insight: Measurement Changes the Measured

When teaching quantum mechanics, Feynman emphasized the fundamental role of observation:

> "The act of observation itself affects the system being observed."

The same applies to AI evaluation:

> "What we choose to measure is what models are incentivized to optimize. Our evaluation criteria are shaping the evolutionary direction of AI systems."

If benchmarks only measure "exam scores," models will optimize "exam scores"—even at the cost of creativity, common sense, or ethical judgment.

6. Takeaways

If you evaluate or use AI systems, ask yourself:

1. "Do my evaluation criteria reflect what I actually care about?" 2. "Does my evaluation consider AI's social impact?" 3. "Have I listened to different stakeholders?" 4. "Is my evaluation promoting the direction of AI development I hope to see?"

AI evaluation is not a neutral scientific activity. It is a social act with far-reaching social consequences.

When we design evaluation criteria, we are not only measuring AI. We are also defining: what counts as a "good" AI? What counts as a "successful" AI? And, ultimately, what kind of society we want to build with AI?

Tags

#ai-evaluation#sociotechnical-systems#generative-ai#pluralism#ai-ethics#benchmarks#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619294