Overview
Research area: Agents Authors: Lei Bai, Zongsheng Cao, Yang Chen Published: 2026-07-01 arXiv: 2507.00010
Abstract
We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajectories and scaling heterogeneous agent abilities.
To support this goal, we build a long-horizon knowledge-action infrastructure that connects external knowledge, actions, observations, and verifier outcomes, producing agentic trajectories with an average length of 45K tokens.
Based on this, we train Agents-A1 with a three-stage recipe:
1. Full-domain supervised fine-tuning to align the base model with broad agentic behaviors. 2. Domain-level teacher models to capture specialized expertise in each domain. 3. Multi-teacher domain-routing on-policy distillation with significant vocabulary alignment to improve cross-domain knowledge-transfer efficiency, unifying six heterogeneous domains into a single deployable student model.
Results
Agents-A1 achieves strong and broad performance on long-horizon agent benchmarks. Compared with 1-trillion-parameter models such as Kimi-K2.6 and DeepSeek-V4-pro, Agents-A1 delivers leading results on:
- SEAL-0: 56.4
- IFBench: 80.6
- HiPhO: 46.4
- FrontierScience-Olympiad: 79.0
- MolBench-Bind: 56.8
- SciCode: 44.3
- HLE: 47.6
- BrowseComp: 75.5
It remains highly competitive on:
--- *Auto-collected on 2026-07-01*