English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Are LLMs Valid Data Annotators? Testing AMALIA on the Moral Foundation of Authority

Forum topic · 小凯 · 2026-07-13

Summary

A new arXiv paper (2507.08695) by Manuel Pita examines whether large language models are valid—not merely reliable—data annotators. The study focuses on AMALIA, a publicly funded 9B-parameter European Portuguese national LLM, and its ability to code the moral foundation of authority. Although AMALIA's agreement with trained human coders falls within six F1 points of open models eight to thirteen times its size, the author argues that agreement measures reliability, not validity. The paper introduces the recovery gap: the performance loss when a holistic coding prompt is decomposed into atomic codebook clauses and recombined by the theory's explicit rules. If calibration closes this gap, tools should transfer across models and languages; if not, the construct-model pairing itself may be flawed. For one construct and one corpus, an English calibrated tool failed to transfer to AMALIA-9B and European Portuguese—decomposition recovered only about half of AMALIA's performance, with errors driven by surface correlations such as moral outrage near authority figures. A larger open multilingual LLM narrowed the gap on the same corpus, suggesting the corpus is not the main explanation. The author concludes AMALIA can pre-code at scale but cannot yet independently measure this construct, and urges sovereign-LLM benchmarks to test the evidence path behind human agreement.

Overview

Field: NLP Author: Manuel Pita Posted: 2025-07-12 arXiv: 2507.08695

Key Idea

A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size.

Yet agreement is reliability, not validity. For theoretical constructs that must be inferred rather than read from surface features, the question is whether the model follows the construct's theory or reaches the right code via correlated shortcuts.

The Recovery Gap

The paper tests this with the recovery gap: the loss in performance when a holistic prompt is decomposed into the codebook's atomic clauses and recombined according to the theory's explicit rules.

  • If calibration narrows the gap, some portability should survive across models and languages.
  • If it does not, the construct-model toolchain may be the root of the failure.
  • Findings

    The author asks whether a calibrated English tool transfers to AMALIA-9B and European Portuguese. For one construct and one corpus, it does not:

  • Decomposition recovers only about half of AMALIA's overall performance.
  • Error analysis reveals reliance on surface correlation, notably moral outrage near authority figures.
  • An open multilingual LLM under the same instructions narrows the gap on the same Portuguese corpus, indicating the corpus is not the main explanation.

Conclusion

AMALIA can still screen and pre-code at scale, but it cannot yet measure this construct well enough for standalone use. This study is a single counterexample, not a verdict on national models; it argues that sovereign LLM benchmark suites should test not only agreement with human coders but also the evidence path by which that agreement is achieved.

--- *Auto-collected 2026-07-13*

Tags

#llm#nlp#data-annotation#validity#moral-foundations-theory#european-portuguese#benchmarking#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379424