Summary
A 2025 arXiv paper (2507.08695) by Manuel Pita examines whether large language models are valid data annotators, using AMALIA, Portugal's publicly funded 9B-parameter European Portuguese language model. On agreement alone, AMALIA looks competitive: when asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. The paper argues that agreement is reliability, not validity, and introduces the recovery gap as a test: the performance loss when a holistic annotation prompt is decomposed into the codebook's atomic clauses and recombined according to the theory's explicit rules. For one construct and one corpus, calibration does not transfer. Decomposition recovers only about half of AMALIA's holistic performance, and error analysis suggests the model relies on surface correlations, particularly moral outrage near authority figures. The findings caution against equating inter-annotator agreement with construct validity in LLM-based annotation pipelines.
Overview
Field: NLP
Author: Manuel Pita
Published: 2025-07-12
arXiv: 2507.08695
Abstract
A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. Yet agreement is reliability, not validity. The paper tests this with the recovery gap: the loss in performance when a holistic prompt is decomposed into the codebook's atomic clauses and recombined by the theory's explicit rule.
Key Findings
- For one construct and one corpus, calibrated performance does not transfer under decomposition.
- Decomposition recovers only about half of AMALIA's holistic performance.
- Error analysis suggests reliance on surface correlations, especially moral outrage near authority figures.
Takeaway
High agreement with human coders demonstrates reliability, not construct validity. LLM-based annotation pipelines should probe whether models follow the underlying coding theory or merely exploit surface cues.
---
*Auto-collected on 2025-07-13*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178379433