[论文] From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for...

研究领域: ML 作者: Stefan G. Creadore 发布时间: 2026-10-05 arXiv: 2610.00015

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Stefan G. Creadore 发布时间: 2026-10-05 arXiv: 2610.00015

中文摘要

大语言模型智能体可以提议并执行动作,但「提议」「授权」「分发」「已验证的外部效果」和「上线推广」是不同的主张。我们提出 Praxa,一种智能体执行框架,通过确定性的准入、代理执行、外部回读、对账和评审推广来显式表示这些状态。四条证据线:第一,作者在固定版本上运行的仓库本地审计通过 1,027/1,027 个单元测试和 89/89 个 Workerd 测试,覆盖全部 363 个预期源文件并满足四项覆盖率下限,但缺少原始逐测试记录和独立复现。第二,提供商支持的 Terminal-Bench Core 0.1.1 试点中,基线和可靠性层各通过 17/36 次严格试验;可靠性层多消耗 37.49% 输入 token 和 50.73% 输出 token,试点不支持优越性结论。第三,调试后的两阶协同代理开发比较中,基线和源码候选各完成 180/180 次试验,准确率相等,完整密封崩溃恢复且零受保护违规;候选 token 减少 37.11%,估计成本降低 33.84%,步骤减少 11.63%,但不能确立质量或延迟改善。第四,部署证据显示有界的反思、召回核算、记忆编译和工具健康路径,但无生产效果提升。Praxa 的可靠贡献是证据绑定的架构,使「授权到效果」的转移显式且可测试。当前证据不能确立对抗安全、生产安全、通用专家优越性、自主递归优化或用户收益。

原文摘要

Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims. We present Praxa, an agent harness that represents these states explicitly through deterministic admission, brokered execution, external read-back, reconciliation, and reviewed promotion. We report four evidence lanes. First, an author-run repository-local audit at a pinned revision passed 1,027/1,027 unit tests and 89/89 Workerd tests, instrumented all 363 expected source files, and met four coverage floors; raw per-test transcripts and independent reproduction are unavailable. Second, in a provider-backed Terminal-Bench Core 0.1.1 pilot across 12 curated tasks, baseline and reliability-layer arms each passed 17/36 strict tri...


*自动采集于 2026-10-05*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens