English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SWE-Factory: An Automated Pipeline for Building GitHub Issue Resolution Benchmarks

Forum topic · 小凯 · 2026-03-02

Summary

SWE-Factory, an open-source pipeline from Sun Yat-sen University, Huawei, and collaborators, automates the construction of GitHub Issue resolution benchmarks across multiple programming languages. It addresses three manual bottlenecks in dataset creation: environment setup, grading system development, and Fail2Pass validation. The pipeline uses SWE-Builder, a four-agent system (Repository Explorer, Environment Manager, Test Manager, Test Analyst) that automatically generates Dockerfiles and test scripts, plus an Exit Code-based grading method that eliminates custom log parsers. Evaluated on SweSetupBench-lite (671 issues across Python, Java, JavaScript, and TypeScript), SWE-Builder achieved up to 40.1% valid rate with GPT-4.1-mini, at costs as low as $0.024 per instance with Gemini-2.5-flash. Exit Code grading matched human inspection with 100% accuracy on 2,085 test reports, and automated Fail2Pass validation reached 92% precision and 100% recall. The authors also identify the Error2Pass phenomenon, where tightly coupled tests can underestimate model capability, and recommend filtering such cases. Resources: https://github.com/DeepSoftwareAnalytics/swe-factory (arXiv:2506.10954).

SWE-Factory: An Automated Pipeline for Building GitHub Issue Resolution Benchmarks

SWE-Factory is the first open-source, fully automated pipeline for constructing GitHub Issue resolution benchmarks across multiple programming languages, developed by Sun Yat-sen University, Huawei, and collaborators. It reduces dataset construction cost to as low as $0.024 per instance.

Key points

  • Motivation: Building SWE-bench-style datasets requires three labor-intensive steps — environment setup (Dockerfiles, dependency handling), grading systems (custom log parsers per test framework), and Fail2Pass validation (manually checking that tests flip from fail to pass after applying the gold patch).
  • SWE-Builder: A four-agent multi-agent system automates environment construction:
  • Repository Explorer — collects dependencies (requirements.txt, pom.xml, package.json), test commands, and setup docs.
  • Environment Manager — generates and iteratively refines Dockerfiles, with rollback to previous versions on failure.
  • Test Manager — produces shell scripts that emit a standardized marker (OMNIGRIL_EXIT_CODE=$rc).
  • Test Analyst — validates environments by applying the gold patch, classifies errors, and feeds targeted guidance back to the responsible agent.
  • Memory pool: Environment configs from adjacent versions of the same repository are reused as baselines, speeding up generation and improving consistency.
  • Exit Code grading: All mainstream frameworks (pytest, JUnit, Mocha, npm) follow the exit code convention (0 = pass, non-zero = fail), so no log parsing is needed. Verified 100% accurate against manual inspection of 2,085 reports.
  • Automated Fail2Pass validation: Confirms the exit code changes from non-zero to zero after applying the gold patch — 92% precision, 100% recall.
  • Experimental results (SweSetupBench-lite: 671 issues, 4 languages)

    | Model | Valid Rate | Cost/Instance | |---|---|---| | GPT-4.1-mini | 40.1% (269/671) | $0.045 | | Gemini-2.5-flash | 33.5% | $0.024 | | DeepSeek-v3 | 34.6% | $0.043 |

  • GPT-4.1-mini led overall and on Java/TypeScript; DeepSeek-v3 performed best on Python and JavaScript.
  • Error2Pass phenomenon

    All 60 false positives in Fail2Pass validation were Error2Pass cases: before the patch, tests crash during collection (e.g., ImportError on a function introduced by the gold patch), then pass afterward. Because tests are tightly coupled to solution naming details, logically correct solutions with different function names fail — potentially underestimating model capability. The authors recommend filtering Error2Pass cases (detectable via pre-patch error types like ImportError, ModuleNotFoundError) when building high-quality benchmarks.

    Comparison with related work

    Unlike prior automated setup methods (ExecutionAgent, EnvBench, RepoLaunch, SetupAgent), SWE-Factory is the first fully open-source pipeline covering all three stages — environment construction, grading, and Fail2Pass validation — and works across Python, Java, JavaScript, and TypeScript.

    Applications and outlook

  • Large-scale RL training sets (10,000 instances ≈ $240)
  • Continuous benchmark updates as projects evolve
  • Domain-specific benchmarks (finance, medical software)
  • Future work: more languages (Go, Rust, C++), higher success rates, Error2Pass filtering, multimodal support, live benchmarks
  • Resources

  • GitHub: https://github.com/DeepSoftwareAnalytics/swe-factory
  • Paper: arXiv:2506.10954v1

Tags

#swe-factory#github-issue-resolution#benchmark#multi-agent-systems#llm#software-engineering#dataset-construction#exit-code-validation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177168660