From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
Overview
- Field: NLP
- Authors: Laurens Samson, Iva Gornishka, Gossa Lô
- Published: 2026-08-11
- arXiv: 2508.03803
- Advisory board process, user research, and a survey of civil-servant chatbot users were used to identify evaluation criteria.
- Six evaluation dimensions were identified: 1. Factuality 2. Honesty 3. Social bias 4. Energy consumption 5. Cost 6. Training data transparency
- These dimensions were operationalised into a benchmark suite covering more than 30 multilingual and Dutch-specific models.
- No single model excels across all dimensions — trade-offs are unavoidable.
- Higher quality consistently comes with greater environmental impact and financial cost.
- Social bias is largely independent of both quality and cost.
- Factuality and honesty are driven by different model properties: a model that answers correctly does not necessarily admit when it does not know.
- A public, user-friendly model overview has been released to make the results actionable for the full range of stakeholders involved in government LLM selection, from engineers to policymakers.
Abstract
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. The authors present the 'Grip on LLMs' framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation.
Methodology
Key Findings
*Originally collected on 2026-08-12.*