Paper Overview
- Field: NLP
- Authors: Laurens Samson, Iva Gornishka, Gossa Lô
- Published: 2026-08-11
- arXiv: 2508.03803
- No single best model: No model excels across all six dimensions, and trade-offs are unavoidable.
- Quality vs. cost trade-off: Higher output quality consistently correlates with greater environmental impact and financial cost.
- Bias independence: Social bias is largely independent of both quality and cost.
- Factuality ≠ Honesty: Factuality (whether a model gives correct answers) and honesty (whether a model acknowledges not knowing) are driven by different model attributes; high factuality does not imply high honesty.
- arXiv: https://arxiv.org/abs/2508.03803
- Auto-collected: 2026-08-12
Summary
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the 'Grip on LLMs' framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation.
Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions:
1. Factuality 2. Honesty 3. Social bias 4. Energy consumption 5. Cost 6. Training data transparency
These dimensions are operationalised into a benchmark suite covering more than 30 multilingual and Dutch-specific models.
Key Findings
Practical Contribution
To make the findings actionable for non-technical audiences, the authors release a publicly accessible, user-friendly model overview tailored to the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.