Paper overview
- Research area: NLP
- Author: Manuel Pita
- Posted: 2025-07-12
- arXiv: 2507.08695
- Construct: Annotation of the authority moral foundation in Portuguese text.
- Reliability vs. validity: High agreement with trained human coders (within six F1 points of much larger open models) demonstrates reliability, not necessarily construct validity.
- Recovery gap method: Holistic prompts are decomposed into atomic codebook clauses and recombined under explicit theoretical rules; the resulting performance loss diagnoses whether the model encodes the theory or exploits correlated shortcuts.
- Main empirical finding: On one construct and one Portuguese corpus, decomposition recovered only about half of AMALIA's holistic performance, indicating shortcut reliance.
- Error pattern: Models latch onto surface cues such as moral outrage in the vicinity of authority figures.
- Corpus vs. model explanation: An open multilingual LLM applied to the same Portuguese corpus narrowed the gap, suggesting the corpus itself is not the primary source of the failure.
- Practical implication: AMALIA is usable for large-scale screening and pre-annotation but not yet sufficient as a standalone measurement instrument for this construct.
- Methodological recommendation: Sovereign-LLM benchmarking suites should evaluate both human-coder agreement and the evidential pathways producing it.
- Caveat: The study is a single counterexample for one model, construct, and corpus, not a general verdict on national-language LLMs.
- arXiv: https://arxiv.org/abs/2507.08695
English abstract
A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. Yet agreement is reliability, not validity. For theoretical constructs that must be inferred rather than read from surface features, the question is whether the model follows the construct's theory or reaches the right code by correlated shortcuts. The authors test this with the recovery gap: the loss in performance when a holistic prompt is decomposed into the codebook's atomic clauses and recombined by the theory's explicit rules. If calibration closes this gap, some portability should survive across models and languages; if not, construct-model tools may be the failure locus. They ask whether a calibrated English tool can transfer to AMALIA-9B and European Portuguese. For one construct and one corpus, it cannot. Decomposition recovered only about half of AMALIA's holistic performance, and error analysis indicated reliance on surface correlations, especially moral outrage near authority figures. An open multilingual LLM narrowed the gap on the same Portuguese corpus under the same instructions, showing the corpus is not the main explanation. AMALIA can still be used for large-scale screening and pre-annotation, but it is not yet good enough to measure this construct independently. The study is a single counterexample, not a verdict on national models. It argues that sovereign LLM benchmarking suites should test not only agreement with human coders, but also the evidential pathways on which that agreement rests.