Overview
Field: NLP Author: Manuel Pita Posted: 2025-07-12 arXiv: 2507.08695
Key Idea
A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size.
Yet agreement is reliability, not validity. For theoretical constructs that must be inferred rather than read from surface features, the question is whether the model follows the construct's theory or reaches the right code via correlated shortcuts.
The Recovery Gap
The paper tests this with the recovery gap: the loss in performance when a holistic prompt is decomposed into the codebook's atomic clauses and recombined according to the theory's explicit rules.
- If calibration narrows the gap, some portability should survive across models and languages.
- If it does not, the construct-model toolchain may be the root of the failure.
- Decomposition recovers only about half of AMALIA's overall performance.
- Error analysis reveals reliance on surface correlation, notably moral outrage near authority figures.
- An open multilingual LLM under the same instructions narrows the gap on the same Portuguese corpus, indicating the corpus is not the main explanation.
Findings
The author asks whether a calibrated English tool transfers to AMALIA-9B and European Portuguese. For one construct and one corpus, it does not:
Conclusion
AMALIA can still screen and pre-code at scale, but it cannot yet measure this construct well enough for standalone use. This study is a single counterexample, not a verdict on national models; it argues that sovereign LLM benchmark suites should test not only agreement with human coders but also the evidence path by which that agreement is achieved.
--- *Auto-collected 2026-07-13*