[论文] Studying Without a Syllabus: Task-Agnostic Environment Preprocessing

论文概要 研究领域: cs.AI, cs.CL, cs.LG 作者: Vinay Samuel, Varun Ursekar, Vijay S. Kalmath, Apaar Shanker, Veronica Chatrath, Yuan Xue 发布时间: 2026-09-13 arXiv: 2609.10824

论文概要

研究领域: cs.AI, cs.CL, cs.LG 作者: Vinay Samuel, Varun Ursekar, Vijay S. Kalmath, Apaar Shanker, Veronica Chatrath, Yuan Xue 发布时间: 2026-09-13 arXiv: 2609.10824

中文摘要

在 LLM 智能体处理新环境中的任务之前,它可以检查可用的语料库和工具,并构建可重用的资源,如索引、脚本或程序指导。然而,大多数自动适应方法依赖任务示例、轨迹或评估反馈来决定构建什么。现有的任务无关方法避免了这种监督,但预先承诺了特定类型环境的准备策略。我们研究更开放的设置:智能体能否在没有教学大纲的情况下学习不熟悉的环境,即在测试时间之前且不知道下游任务分布的情况下选择如何准备?我们将任务无关环境预处理形式化,其中学习系统在预算下探索环境并为冻结的求解器生成工件。我们在六个异构基准上比较无辅助和档案装备的元智能体与固定合成实践和语料库处理方法。元智能体变体在五个基准上实现最高的 Avg@3 奖励,而固定语料库处理在最大语料库基准上保持最佳。

原文摘要

Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusable resources such as indices, scripts, or procedural guidance. Most automated adaptation methods, however, rely on task examples, trajectories, or evaluation feedback to decide what to build. Existing task-agnostic approaches avoid this supervision but commit in advance to a preparation strategy for a particular type of environment. We study a more open-ended setting: can an agent study an unfamiliar environment without a syllabus, i.e. before test time and without knowledge of the downstream task distribution, and choose how to prepare it? We formalize task-agnostic environment preprocessing, in which a studying system explores an environment under a budget and produces artifacts for a frozen solver. We compare unaided and archive-equipped meta-agents with fixed synthetic-practice and corpus-processing methods across six heterogeneous benchmarks. A meta-agent variant achieves the highest Avg@3 reward on five benchmarks, while fixed corpus processing remains best on the largest corpus benchmark. Larger study budgets do not reliably improve downstream reward. Nevertheless, studied artifacts reduce the test-time sampling needed to reach a given score, demonstrating how reusable preparation can shift computation from repeated test-time attempts to a pre-task study phase.


*自动采集于 2026-09-13*

#论文 #arXiv #AI #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens