Paper Overview
Field: NLP Authors: Laurens Samson, Iva Gornishka, Gossa Lô Published: 2026-08-12 arXiv: 2508.05157
Abstract
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. This paper presents the 'Grip on LLMs' framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation.
Methodology
Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, the authors identify six evaluation dimensions:
1. Factuality 2. Honesty 3. Social bias 4. Energy consumption 5. Cost 6. Training data transparency
These dimensions are operationalised into a benchmark suite covering more than 30 multilingual and Dutch-specific models.
Key Findings
- No single model excels across all dimensions — trade-offs are inevitable. Higher quality generally implies greater environmental impact and financial cost, while social bias appears largely unrelated to both.
- Factuality and honesty are distinct: factuality (whether the model answers correctly) and honesty (whether the model admits it does not know) are governed by different attributes; high factuality does not imply high honesty.
- To make the findings accessible to non-technical audiences, the authors released a publicly available, user-friendly model overview tool serving stakeholders from engineers to policymakers.
*Auto-collected on 2026-08-12*