[论文] Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM ...
研究领域: ML 作者: Jundong Hu, Shekar Ramachandran 发布时间: 2026-10-05 arXiv: 2610.00025
论文概要
研究领域: ML 作者: Jundong Hu, Shekar Ramachandran 发布时间: 2026-10-05 arXiv: 2610.00025
中文摘要
智能体执行框架越来越希望在前沿大模型规划器周围的微任务上运行小语言模型(SLM):自动批准 shell 命令、写记忆、选工具、给历史轮次排序。我们问:现成 SLM 是否达到从业者定义的阈值?失败原因是什么?量化会否改变答案?我们构建含 4 个此类微任务的基准,固定提示和自动指标,每个任务有预指定阈值 τ,锚定廉价非 LLM 基线和 CI 感知资格规则(配置的置信界超过 τ 才算通过)。在 Qwen3 0.6/1.7/4/8B 最优设置下扫描(FP16、贪心、单一冻结提示、无调优),发现资格缺口:16 个(4 任务 × 4 模型)配置中 0 个通过(已检查原始输出和解析器行为验证)。Logprob 决策阈值诊断将失败分为能力缺陷和可经解码阈值解决的失败(4 种状态)。4 比特量化(RTN/GPTQ/AWQ)的损害取决于模型大小,未使任何配置合格;缺口更多跟随规模而非精度。Llama-3.x 上可复现(12/12 不合格),对锚点选择和提示措辞稳健(含 3 个中性改写共 0/112 合格)。实践含义:将 SLM 置于可达 CI 阈值的基线之后,仅在基线不达标处使用;如 4B 重排序器在 BM25 短名单上优于 BM25(+0.047 [0.020, 0.073],但自身仍未认证合格)。
原文摘要
Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell commands, writing memory, selecting tools, ranking past turns. We ask whether off-the-shelf SLMs meet practitioner-defined thresholds and, when they fail, why, and whether quantization changes the answer. We build a benchmark of 4 such microtasks with fixed prompts and automatic metrics, each with a pre-specified threshold \(\tau\) anchored to a cheap non-LLM baseline and a CI-aware eligibility rule (a configuration passes only if its confidence bound clears \(\tau\)). Sweeping Qwen3 0.6/1.7/4/8B at their best (FP16, greedy, one frozen prompt, no tuning), we find an eligibility gap: 0 of 16 (4 tasks \(\times\) 4 models) configurations ...
*自动采集于 2026-10-05*
#论文 #arXiv #ML #小凯