English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Laguna M.1/XS.2 Technical Report: MoE Foundation Models for Agentic Coding

Forum topic · 小凯 · 2026-05-29

Summary

This arXiv paper (2605.27605) introduces Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding. M.1 has 225.8B total parameters with 23.4B activated per token, while XS.2 has 33.4B total parameters with 3B activated. Both models were trained from scratch end-to-end inside an internal system called the Model Factory—a tightly integrated, versioned stack of data, training, evaluation, and inference components that turns model development into an industrial process. The report details the Model Factory's design principles and the full training pipeline, covering pre-training data and architecture, post-training stages, evaluation, and quantization. On agentic software engineering and terminal benchmarks including SWE-bench Verified, SWE-bench Multilingual, SWE-Bench Pro, and Terminal-Bench 2.0, both models are competitive with state-of-the-art open-weight models in their weight classes. Laguna XS.2 weights have been open-sourced under the Apache 2.0 license on Hugging Face.

Overview

Field: AI Authors: Julien Abadji, Marah Abdin, Connor Adams, et al. Published: 2026-05-28 arXiv: 2605.27605

Abstract (translated)

This paper presents Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has 225.8B total parameters (23.4B activated per token), and XS.2 has 33.4B total parameters (3B activated per token). Both models were trained from scratch end-to-end inside an internal system called the "Model Factory"—a tightly integrated stack of versioned data, training, evaluation, and inference components that turns model development into an industrial process.

The paper details the design principles and choices of the Model Factory, as well as the end-to-end training process of the models, covering pre-training data and architecture, post-training stages, evaluation, and quantization.

On agentic software engineering and terminal benchmarks (SWE-bench Verified, SWE-bench Multilingual, SWE-Bench Pro, and Terminal-Bench 2.0), M.1 and XS.2 are competitive with state-of-the-art open-weight models in their respective weight classes. The Laguna XS.2 weights have been open-sourced under the Apache 2.0 license on Hugging Face.

Original Abstract (excerpt)

We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has 225.8B total parameters (23.4B activated per token) and XS.2 has 33.4B total (3B activated). Both models were trained from scratch end-to-end inside the same internal system that we refer to as our Model Factory: a tightly-integrated stack of versioned data, training, evaluation, and inference components that turn model development into an industrial process. We describe the principles and design choices of the Model Factory and also detail the end-to-end training process of our models, throughout pre-training data and architecture, post-training stages, evaluation, and quantization. On agentic software engineering and terminal benchmarks (SWE-bench Verified, SWE-bench Multilingual, SWE-Bench Pro, and Terminal-Bench 2.0) M.1 and XS.2 are competitive with state-o...

--- *Auto-collected on 2026-05-29*

Tags

#ai#mixture-of-experts#foundation-models#agentic-coding#swe-bench#open-source#arxiv#model-training

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980491