Paper Overview
Field: NLP (Natural Language Processing) Authors: Laurens Samson, Iva Gornishka, Gossa Lô Published: 2026-08-11 arXiv: 2508.03803
Summary
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. The authors present the 'Grip on LLMs' framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation.
Through an advisory board process, user research, and a survey of civil-servant chatbot users, the paper identifies six evaluation dimensions:
- Factuality
- Honesty
- Social bias
- Energy consumption
- Cost
- Training data transparency
- No single model excels across all dimensions; trade-offs are unavoidable.
- Higher model quality consistently correlates with greater environmental impact and financial cost, while social bias appears largely independent of both.
- Factuality and honesty are driven by different model properties — high factuality does not imply high honesty. Factuality measures whether a model answers correctly, whereas honesty measures whether a model acknowledges when it does not know.
These dimensions are operationalised into a benchmark suite covering more than 30 multilingual and Dutch-specific models.
Key Findings
Practical Contribution
To make these findings actionable for non-technical audiences, the authors release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection — from engineers to policymakers.
Source
*Auto-collected from arXiv on 2026-08-12.*
#arXiv #NLP #LLM-Evaluation #Dutch-Government #Benchmark #Public-Administration