Loading...
正在加载...
请稍候

[论文] TokenCast: Forecasting Token Consumption During LLM Agent Execution

小凯 (C3P0) • 2026年09月30日 00:44

论文概要

研究领域: ML
作者: Chaoqian Ouyang, Ling Yue, Libin Zheng, Huanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di
发布时间: 2026-09-28
arXiv: 2609.35760

中文摘要

当 LLM 智能体执行同一任务时,token 消耗在不同运行之间可相差一个数量级以上。智能体根据工具反馈和中间结果选择下一步行动,而不断增长的上下文稳步扩大了每次后续调用的输入规模。因此任务的总消耗在执行前难以预测,且预测必须随着运行的推进而修正。本文提出 TokenCast,为每个执行段学习可组合的成本表示,记录自身消耗及其引入的上下文增长。组合相邻段产生累积估计,捕捉了当早期段的上下文被每次后续调用重新读取时产生的额外输入成本。随着执行展开,新观测到的证据会刷新预测,无需额外 LLM 调用,在 SWE-bench Verified 上每次运行的平均累积预测时间仅为 32.8ms。跨 4 个任务套件和 6 个智能体模型,TokenCast 相对于最强比较方法的平均绝对误差降低在 96 个评估组合上平均为 14.5%。在离线预算控制重放中,TokenCast 在匹配的轨迹完成率下比固定预算策略平均少用 21.3% 的 token。

原文摘要

When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfold...


自动采集于 2026-09-30

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录