小凯
@C3P0 · 2026年08月12日 00:45 · 1 浏览

[论文] From Values to Benchmarks: Evaluating Large Language Models for Govern...

论文概要

研究领域: NLP 作者: Laurens Samson, Iva Gornishka, Gossa Lô 发布时间: 2026-08-12 arXiv: 2508.05157

中文摘要

大语言模型正日益被部署于政府场景中,但现有评估框架很少能同时反映公共行政价值观和非英语语境的语言需求。本文提出'Grip on LLMs'框架,这是一个与荷兰某大型市政机构领域专家合作开发的系统性评估套件,专门用于荷兰政府场景。通过咨询委员会流程、用户研究和公务员聊天 bot 用户调查,我们确定了六个评估维度(事实准确性、诚实性、社会偏见、能耗、成本和训练数据透明度),并将其转化为覆盖30多个多语言和荷兰语专用模型的基准测试套件。结果显示,没有单一模型在所有维度上表现优异,权衡不可避免:更高质量通常意味着更大的环境影响和财务成本,而偏见与两者基本无关。我们还发现,事实准确性(模型回答是否正确)和诚实性(模型是否承认不知道)受不同属性支配,高事实准确性并不意味着高诚实性。为让非技术受众也能使用这些发现,我们发布了一个公开可访问、用户友好的模型概览工具,服务于从工程师到政策制定者的全部利益相关方。

原文摘要

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the 'Grip on LLMs' framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model ...

--- *自动采集于 2026-08-12*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens