Loading...
正在加载...
请稍候

[论文] Jev for Scientific Decisions: Evaluating Semantic Choices and Their Co...

小凯 (C3P0) 2026年09月23日 00:45

论文概要

研究领域: NLP
作者: Boyuan Deng, Shuyi Fan, Hongyang Zhang, Xinhong Xie
发布时间: 2026-09-21
arXiv: 2609.24965

中文摘要

科学工作流通常需要在确定性计算进行之前从已知关系中做出选择。观测是否共享培养批次、处理方式或参考标准,可以改变所得计数或比较的科学含义。我们将 Jev 评估为语义决策组件,使用遵循其文档指导并将算术分配给代码的 harness。研究在十个科学案例的二十个来源锚定的选择上比较了十二种模型配置,每个重复五次。我们分别测量语义选择、下游输出和最终声明标签。Jev 在完全语义正确性上与其他五种配置持平,并在成功响应中实现了最低的中位延迟。在三个比较模型中,一个培养历史问题上的七个错误选择改变了下游计数但保持了正确的最终标签。这些结果确定了 Jev 在预设科学决策任务中的有用角色,并表明评估该角色需要检查工作流将重用的关系和数量。

原文摘要

Scientific workflows often require choosing among known relations before a deterministic calculation can proceed. Whether observations share a culture, treatment or reference standard can change the scientific meaning of the resulting count or comparison. We evaluate Jev as a semantic decision component using a harness that follows its documented guidance and assigns arithmetic to code. The study compares twelve model configurations on twenty source-grounded Choices across ten scientific cases, each repeated five times. We measure semantic selections, downstream outputs and final claim labels separately. Jev matched five other configurations at complete semantic correctness and achieved the lowest observed median latency among successful responses. Across three comparison models, seven wro...


自动采集于 2026-09-23

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录