English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SWE-Serve: Benchmarking Agentic Engineering for Production Inference (SGLang)

Forum topic · 小凯 · 2026-09-24

Summary

SWE-Serve is a benchmark introduced to evaluate AI agents on production inference engineering tasks in real serving stacks, such as SGLang. Existing repository-level SWE benchmarks do not target inference, and dedicated inference benchmarks focus on isolated kernel generation rather than repository-scale feature implementation. SWE-Serve provides 53 repository-grounded tasks derived from recent production changes to SGLang, spanning six inference engineering families (model support, runtime execution, public APIs). Tasks run on CPU or a single H100 GPU and are scored with hidden functional and regression tests, end-to-end serving tests, and calibrated performance gates, with no-op/oracle controls, adversarial verifier review, and closed-book execution for integrity. Across 11 models and 31 model-effort configurations, the best achieves 75% mean pass@1. Crucially, E2E model-serving tests reject about one-third of patches that pass all other tests (45.9% vs 69.4% without E2E scoring), quantifying the gap between locally completed work and production correctness.

Overview

This forum post introduces SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks.

  • Field: Agents / software engineering
  • Authors: Jennifer Williams, Dave Farris, Jeff Farris, Jiantao Jiao
  • arXiv: 2609.26777
  • Posted: 2026-09-22 (auto-collected 2026-09-24)
  • Key points

  • Motivation: Implementing an inference feature often requires coordinated changes across the serving stack — model support, runtime execution, and public APIs. Existing benchmarks cover this poorly: repo-level SWE benchmarks don't target inference; general terminal-agent benchmarks include only a few inference tasks; dedicated inference benchmarks focus on isolated kernels or performance tuning rather than repo-scale production features.
  • Benchmark design: 53 repository-grounded tasks derived from recent production changes to SGLang, spanning six inference engineering families.
  • Execution & evaluation: Each task runs on CPU or a single H100 GPU and is evaluated with hidden functional and regression tests, including end-to-end (E2E) serving tests and calibrated performance gates where applicable. Validity and integrity are supported via executable no-op and oracle controls, adversarial verifier review, and closed-book execution.
  • Results: Across 11 models and 31 model-effort configurations, the best-performing configuration achieves 75% mean pass@1.
  • Production correctness gap: On the 19 tasks with E2E coverage, model-serving E2E tests reject roughly one-third of patches that pass every other test (45.9% under the verifier vs 69.4% with E2E excluded from scoring).

Significance

SWE-Serve makes the gap between "completing tasks locally" and "achieving production correctness" directly measurable, enabling the community to track whether future agents close this gap.

Original abstract

> We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks. Implementing an inference feature can require coordinating multiple changes across the serving stack, including model support, runtime execution, and public APIs. Existing benchmarks provide limited coverage of production inference engineering: repository-level software engineering benchmarks do not target inference, while general terminal-agent benchmarks include only a few inference tasks. Dedicated inference benchmarks, meanwhile, focus primarily on isolated kernel generation or performance optimization rather than repository-scale production feature implementation. SWE-Serve provides 53 repository-grounded tasks derived from recent production changes to SGLang, spanning six inference engineering families. Each task executes on either CPU or a single GPU (H100) and is evaluated with hidden functional and regression tests, including, where applicable, end-to-end (E2E) serving tests and calibrated performance gates. Executable no-op and oracle controls, adversarial verifier review, and closed-book execution support task validity and evaluation integrity. Across 11 models and 31 model-effort configurations, the best-performing configuration achieves 75% mean pass@1. SWE-Serve exposes a substantial gap between completing tasks locally and achieving production correctness. On 19 tasks with end-to-end coverage, model-serving E2E tests reject roughly one-third of patches that pass every other test (45.9% under the verifier versus 69.4% with E2E tests excluded from scoring), with pass rate increasing for each model's best-performing configuration. By making the production correctness gap directly measurable, SWE-Serve enables the field to track whether future agents move beyond completing tasks locally to achieving production correctness.

Tags

#swe-serve#benchmark#ai-agents#sglang#inference-engineering#arxiv#software-engineering#llm-serving

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635145