[论文] Policy-as-Skill: Governed LLM Decision Support with Evidence, Determin...

研究领域: ML 作者: Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya 发布时间: 2026-09-25 arXiv: 2609.27087

论文概要

研究领域: ML 作者: Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya 发布时间: 2026-09-25 arXiv: 2609.27087

中文摘要

组织越来越多地使用 LLM 进行政策、合规、风险与运营决策支持,这需要证据验证、审核路由、版本控制与可审计性。我们提出 Policy-as-Skill(PaS):一种模块化运行时,把这些功能打包为可执行、可版本化的策略能力。以固定 Gemma4 后端在 600 个开发任务上评估 13 种方法。PaS+Audit 达到 53.8% 精确准确率、macro-F1 0.346、审核 F1 0.854、引用精确率 1.000、策略引用召回率 0.984、审计完整性 1.000,在多数治理与审核指标上优于 LLM+RAG。确定性控制将总体准确率提升至 61.2%,但强任务依赖,支持选择性而非普适的规则化干预。

原文摘要

Organizations increasingly use LLMs for policy, compliance, risk, and operational decision support, requiring evidence validation, review routing, version control, and auditability. We introduce Policy-as-Skill (PaS), a modular runtime that packages these functions as executable, versioned policy capabilities. Thirteen methods are evaluated with a fixed Gemma4 backend on 600 development tasks. PaS+Audit achieves 53.8% exact accuracy, macro-F1 0.346, review F1 0.854, citation precision 1.000, policy-reference recall 0.984, and audit completeness 1.000, outperforming LLM+RAG on most governance and review metrics. Deterministic control raises aggregate accuracy to 61.2% but is strongly task dependent, supporting selective rather than universal rule-based intervention.


*自动采集于 2026-09-25*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens