[论文] Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, ...
研究领域: ML 作者: Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri, Seong Joon Oh 发布时间: 2026-10-06 arXiv: 2610.08775
论文概要
研究领域: ML 作者: Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri, Seong Joon Oh 发布时间: 2026-10-06 arXiv: 2610.08775
中文摘要
大语言模型(LLM)可以解决许多狭义任务,但对于数百万个相关实例单独查询它们可能成本高得令人望而却步。LLM智能体能否自主为这类工作负载创建更便宜的解决方案?我们将这种能力称为「封装」(bottling):将通用能力转化为平衡答案质量和摊销成本的任务特定解决方案的能力。我们引入BOTTLED基准,智能体在其中接收一个完整的未标注工作负载,必须在固定的时间、计算和LLM API预算下完成它。在十个模型和三个任务中,我们发现强零样本任务表现并不能可靠地转化为强封装能力。零样本得分相似的模型在封装后可能有很大差异,60个封装运行中有48个得分低于其模型零样本表现95%置信区间的下界。尽管如此,封装可以产生可观的节省:在查询-产品相关性分类任务上,Opus 5在约657倍更低的报告成本下保持了约82%的零样本macro-F1。封装也与专为廉价重复推理而建的「系统一」模型Jev具有竞争力。BOTTLED为评估和改进智能体将有限资源投资于大型重复工作负载可重用解决方案的能力提供了基础。
原文摘要
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workloads? We call this ability "bottling": the ability to turn general capabilities into task-specific solutions that balance answer quality and amortised cost. We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets. Agents choose their own approach, such as training a small model or writing a reusable program. Across ten models and three tasks, we find that strong zero-shot task performance does not reliably translate into strong bottling capabilities. Models with similar zero-s...
*自动采集于 2026-10-08*
#论文 #arXiv #ML #小凯