Paper Overview
- Field: NLP
- Authors: Laurens Samson, Iva Gornishka, Gossa Lô
- Published: 2026-08-11
- arXiv: 2508.03803
- No single model excels across all dimensions — trade-offs are unavoidable. Higher quality consistently comes with greater environmental impact and financial cost, while social bias is largely independent of both.
- Factuality vs. honesty: whether a model answers correctly and whether it admits it does not know are driven by different model attributes. High factuality does not imply high honesty.
- Practical output: the team released a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in government LLM selection, from engineers to policymakers.
Summary
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. The authors present the 'Grip on LLMs' framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation.
Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, the study identifies six evaluation dimensions:
1. Factuality 2. Honesty 3. Social bias 4. Energy consumption 5. Cost 6. Training data transparency
These dimensions are operationalized into a benchmark suite covering more than 30 multilingual and Dutch-specific models.
Key Findings
Original Abstract (excerpt)
> Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the 'Grip on LLMs' framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation.
--- *Auto-collected on 2026-08-12*