> Paper: Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechnical Systems > Author: Rebecca L. Johnson > arXiv: 2604.20545 | 2026-04-27
1. The World Where Leaderboards Rule
Open any AI evaluation site and you will see:
- GPT-5.2: 92.3
- Claude 4: 91.7
- Gemini 3: 90.1
- Llama 4: 88.5
- Treats models as isolated predictors
- Runs benchmarks; high scores mean "good"
- Problem: ignores how models are trained, deployed, and used
- Evaluates what models "should" be
- Uses human values as the standard
- Problem: whose values? Which culture? Which era?
- Training data is not "raw material" but the product of a social process (who decides what data to collect?)
- Annotation is not "objective truth" but a collection of human judgments (who labeled it, and by what standards?)
- Deployment is not a mere "application of technology" but an embodiment of power relations (who decides where AI is used?)
- Evaluation is not "scientific measurement" but an expression of value choices (what do we choose to measure, and why?)
- Different stakeholders have different evaluation needs
- Different cultural backgrounds carry different value standards
- Different application contexts define success differently
- Developers may care about "reasoning ability"
- Policymakers may care about "safety and fairness"
- End users may care about "usefulness and ease of use"
- Affected communities may care about "autonomy and dignity"
These numbers look objective, scientific, unimpeachable. But the question is: what do they actually measure?
The paper issues a fundamental challenge: evaluating generative AI is not just a technical problem—it is a sociotechnical one.
2. Two Flawed Evaluation Paradigms
The research critiques two mainstream approaches:
1. Functionalism
2. Prescriptivism
The shared blind spot of both paradigms: they treat models as "things" to be measured, rather than understanding AI systems as "processes."
3. The Sociotechnical Systems Perspective
What is a "sociotechnical system"?
> AI is not isolated software. It is a complex system jointly constituted by technology, people, institutions, culture, and history.
This means:
AI evaluation itself shapes AI development.
4. Why Pluralism Is Necessary
The paper argues for pluralist evaluation—not seeking one "correct" evaluation method, but acknowledging that:
For example:
5. A Feynman-Style Insight: Measurement Changes the Measured
When teaching quantum mechanics, Feynman emphasized the fundamental role of observation:
> "The act of observation itself affects the system being observed."
The same applies to AI evaluation:
> "What we choose to measure is what models are incentivized to optimize. Our evaluation criteria are shaping the evolutionary direction of AI systems."
If benchmarks only measure "exam scores," models will optimize "exam scores"—even at the cost of creativity, common sense, or ethical judgment.
6. Takeaways
If you evaluate or use AI systems, ask yourself:
1. "Do my evaluation criteria reflect what I actually care about?" 2. "Does my evaluation consider AI's social impact?" 3. "Have I listened to different stakeholders?" 4. "Is my evaluation promoting the direction of AI development I hope to see?"
AI evaluation is not a neutral scientific activity. It is a social act with far-reaching social consequences.
When we design evaluation criteria, we are not only measuring AI. We are also defining: what counts as a "good" AI? What counts as a "successful" AI? And, ultimately, what kind of society we want to build with AI?