English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Agents-A1: A 35B MoE Agentic Model Reaching Trillion-Parameter-Level Performance via Agent-Horizon Scaling

Forum topic · 小凯 · 2026-07-01

Summary

Agents-A1 is a 35-billion-parameter Mixture-of-Experts agentic model that achieves trillion-parameter-level performance by scaling the agent horizon rather than model size. The authors scale along two axes: long-horizon trajectories and heterogeneous agent abilities. They build a long-horizon knowledge-action infrastructure linking external knowledge, actions, observations, and verifier outcomes, yielding agentic trajectories averaging 45K tokens. Training follows a three-stage recipe: full-domain supervised fine-tuning, domain-level teacher models, and multi-teacher domain-routing on-policy distillation with vocabulary alignment that unifies six heterogeneous domains into one deployable student model. On long-horizon agent benchmarks, Agents-A1 leads trillion-parameter models such as Kimi-K2.6 and DeepSeek-V4-pro on SEAL-0 (56.4), IFBench (80.6), HiPhO (46.4), FrontierScience-Olympiad (79.0), and MolBench-Bind (56.8), while remaining competitive on SciCode (44.3), HLE (47.6), and BrowseComp (75.5). Paper: arXiv 2507.00010.

Overview

Research area: Agents Authors: Lei Bai, Zongsheng Cao, Yang Chen Published: 2026-07-01 arXiv: 2507.00010

Abstract

We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajectories and scaling heterogeneous agent abilities.

To support this goal, we build a long-horizon knowledge-action infrastructure that connects external knowledge, actions, observations, and verifier outcomes, producing agentic trajectories with an average length of 45K tokens.

Based on this, we train Agents-A1 with a three-stage recipe:

1. Full-domain supervised fine-tuning to align the base model with broad agentic behaviors. 2. Domain-level teacher models to capture specialized expertise in each domain. 3. Multi-teacher domain-routing on-policy distillation with significant vocabulary alignment to improve cross-domain knowledge-transfer efficiency, unifying six heterogeneous domains into a single deployable student model.

Results

Agents-A1 achieves strong and broad performance on long-horizon agent benchmarks. Compared with 1-trillion-parameter models such as Kimi-K2.6 and DeepSeek-V4-pro, Agents-A1 delivers leading results on:

  • SEAL-0: 56.4
  • IFBench: 80.6
  • HiPhO: 46.4
  • FrontierScience-Olympiad: 79.0
  • MolBench-Bind: 56.8
  • It remains highly competitive on:

  • SciCode: 44.3
  • HLE: 47.6
  • BrowseComp: 75.5
We hope this work provides the community with a practical path to match the performance of 1T-parameter models on long-horizon tasks using a 35B agentic model with scaled horizon.

--- *Auto-collected on 2026-07-01*

Tags

#agents#mixture-of-experts#distillation#long-horizon-tasks#llm#paper#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208345