SWE-Factory: An Automated Pipeline for Building GitHub Issue Resolution Benchmarks
SWE-Factory is the first open-source, fully automated pipeline for constructing GitHub Issue resolution benchmarks across multiple programming languages, developed by Sun Yat-sen University, Huawei, and collaborators. It reduces dataset construction cost to as low as $0.024 per instance.
Key points
- Motivation: Building SWE-bench-style datasets requires three labor-intensive steps — environment setup (Dockerfiles, dependency handling), grading systems (custom log parsers per test framework), and Fail2Pass validation (manually checking that tests flip from fail to pass after applying the gold patch).
- SWE-Builder: A four-agent multi-agent system automates environment construction:
- Repository Explorer — collects dependencies (
requirements.txt,pom.xml,package.json), test commands, and setup docs. - Environment Manager — generates and iteratively refines Dockerfiles, with rollback to previous versions on failure.
- Test Manager — produces shell scripts that emit a standardized marker (
OMNIGRIL_EXIT_CODE=$rc). - Test Analyst — validates environments by applying the gold patch, classifies errors, and feeds targeted guidance back to the responsible agent.
- Memory pool: Environment configs from adjacent versions of the same repository are reused as baselines, speeding up generation and improving consistency.
- Exit Code grading: All mainstream frameworks (pytest, JUnit, Mocha, npm) follow the exit code convention (0 = pass, non-zero = fail), so no log parsing is needed. Verified 100% accurate against manual inspection of 2,085 reports.
- Automated Fail2Pass validation: Confirms the exit code changes from non-zero to zero after applying the gold patch — 92% precision, 100% recall.
- GPT-4.1-mini led overall and on Java/TypeScript; DeepSeek-v3 performed best on Python and JavaScript.
- Large-scale RL training sets (10,000 instances ≈ $240)
- Continuous benchmark updates as projects evolve
- Domain-specific benchmarks (finance, medical software)
- Future work: more languages (Go, Rust, C++), higher success rates, Error2Pass filtering, multimodal support, live benchmarks
- GitHub: https://github.com/DeepSoftwareAnalytics/swe-factory
- Paper: arXiv:2506.10954v1
Experimental results (SweSetupBench-lite: 671 issues, 4 languages)
| Model | Valid Rate | Cost/Instance | |---|---|---| | GPT-4.1-mini | 40.1% (269/671) | $0.045 | | Gemini-2.5-flash | 33.5% | $0.024 | | DeepSeek-v3 | 34.6% | $0.043 |
Error2Pass phenomenon
All 60 false positives in Fail2Pass validation were Error2Pass cases: before the patch, tests crash during collection (e.g., ImportError on a function introduced by the gold patch), then pass afterward. Because tests are tightly coupled to solution naming details, logically correct solutions with different function names fail — potentially underestimating model capability. The authors recommend filtering Error2Pass cases (detectable via pre-patch error types like ImportError, ModuleNotFoundError) when building high-quality benchmarks.
Comparison with related work
Unlike prior automated setup methods (ExecutionAgent, EnvBench, RepoLaunch, SetupAgent), SWE-Factory is the first fully open-source pipeline covering all three stages — environment construction, grading, and Fail2Pass validation — and works across Python, Java, JavaScript, and TypeScript.