Python Is Not a Proxy: Multi-LCB Exposes the Python Overfitting of Code LLMs
The GigaCode team extended LiveCodeBench (LCB) from Python-only into Multi-LCB, covering 12 programming languages, and evaluated 24 mainstream large models. The result is striking: almost all models are "Python-specialized" — dominant in Python, but much weaker elsewhere. GPT-OSS-120B is the only model with relatively balanced cross-language performance, while Qwen3-235B, though strongest in Python, is overtaken by GPT-OSS-120B on Go, Rust, and Ruby.
Background: What Does LiveCodeBench Miss?
LiveCodeBench fixed three major flaws of earlier benchmarks (HumanEval, MBPP):
1. Data contamination: date filtering ensures only problems published after a model's training cutoff are tested 2. Benchmark saturation: continuous scraping of new problems from LeetCode, AtCoder, and Codeforces 3. Narrow scope: tests full problem solving, not just function completion
But LCB has one fatal blind spot: it only tests Python. Multi-LCB asks a long-avoided question: is LLM coding ability "general programming intelligence" or just "Python memorization"?
Benchmark Design: 12 Languages, 24 Models, One Problem Set
The 12 languages were chosen by cross-referencing four 2025 rankings (TIOBE, GitHub Octoverse, Stack Overflow, RedMonk): Python, C++, Java, C#, Go, Rust, JavaScript, TypeScript, Ruby, PHP, Kotlin, and Scala.
The key technical innovation is an automatic conversion pipeline that transforms LeetCode's function-style problems into unified STDIN/STDOUT format:
- Scalar inputs/outputs: direct mapping
- 1-D arrays: space-separated values
- 2-D arrays: first line gives row count, following lines give space-separated values
- Python is not a proxy: Qwen3-235B leads Python (74.0%) but is beaten by GPT-OSS-120B on Go (56.7%), Rust (47.7%), and Ruby (49.4%). Choosing a model by Python rank alone can mislead.
- GPT-OSS-120B's consistency is remarkable: standard deviation across 12 languages is only 5.9%, versus 9.4% for Qwen3-235B.
- Language difficulty tiers: Python averages ~48.2% Pass@1, Java/C++ ~44%, most other languages ~33–39%, Scala <29%. Python pass rates run 10–20 points higher on the same problem set.
- Python overfitting: nearly all models score systematically higher on Python than their cross-language average. OpenReasoning-Nemotron-32B and OpenCodeReasoning-Nemotron-1.1-32B score 64%+ on Python but under 30% on other languages — a gap exceeding 30 points.
- Statically typed languages are harder than dynamically typed ones, likely due to training-data skew, stricter syntax constraints, longer reasoning chains, and delayed (compile-time) error feedback.
- Benchmark rankings based on Python-only tests (HumanEval, MBPP, LCB) likely carry systematic bias; top-ranked models may simply be better at Python.
- The authors suggest performance could improve by increasing training coverage of non-Python languages.
- Practical model selection: multi-language/full-stack teams should favor GPT-OSS-120B; Python-centric teams may prefer Qwen3-235B-Thk or DeepSeek-R1; systems programming teams should check C++/Rust-specific scores.
- Paper: *Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages*
- Authors: Maria Ivanova, Pavel Zadorozhny, Rodion Levichev, Ivan Petrov, Pavel Adamenko, Ivan Lopatin, Alexey Kutalev, Dmitrii Babaev (GigaCode, Yandex School of Data Analysis)
- Paper: https://arxiv.org/abs/2606.20517
- Code: https://github.com/Multi-LCB/Multi-LCB
- Submitted to ICLR 2026
This preserves algorithmic difficulty while enabling uniform cross-language evaluation.
Key Results
Top models (Feb–May 2025 data, Pass@1)
| Model | Python | C++ | Java | Go | Rust | Ruby | Scala | Cross-lang avg | |-------|--------|-----|------|----|------|------|-------|----------------| | GPT-OSS-120B* (Medium) | 71.1 | 72.3 | 70.4 | 69.9 | 70.5 | 70.2 | 54.1 | 67.8 | | Qwen3-235B-A22B-Thk* | 74.0 | 75.8 | 73.9 | 56.7 | 47.7 | 49.4 | 57.6 | 64.0 | | DeepSeek-R1-0528* | 66.3 | 68.0 | 67.8 | 55.0 | 63.1 | 62.4 | 62.3 | 63.1 | | GPT-OSS-20B* (Medium) | 63.6 | 65.7 | 62.7 | 59.9 | 61.9 | 61.7 | 43.1 | 59.8 | | Qwen3-30B-A3B-Thk* | 64.0 | 65.7 | 62.4 | 44.1 | 51.7 | 42.1 | 43.6 | 53.2 |
(* = reasoning models)
Main findings
Contamination varies by language
Scores drop sharply after models' training cutoffs, but the drop magnitude differs across languages — suggesting asymmetric, language-dependent leakage that date filtering alone does not eliminate.
Reasoning vs. instruction models
| Type | Model | Python | Cross-lang avg | |------|-------|--------|----------------| | Reasoning | GPT-OSS-120B* | 71.1% | 67.8% | | Reasoning | Qwen3-235B-Thk* | 74.0% | 64.0% | | Reasoning | DeepSeek-R1* | 66.3% | 63.1% | | Instruction | Qwen3-235B-Instr | 43.8% | 37.5% | | Instruction | Qwen2.5-Coder-32B | 27.5% | 25.0% |
Reasoning models outperform by a qualitative margin. Notably, DeepSeek-R1-Distill models show extreme language divergence (e.g., Distill-Qwen-14B: Python 41.8%, Rust 3.7%, Scala 3.3%) — distillation appears to amplify the Python bias.
Implications
Limitations
1. Missing languages: Swift, Haskell, R, Julia 2. Only competitive-programming tasks — no API integration, debugging, or code review 3. Forced STDIN/STDOUT conversion may disadvantage some languages 4. Closed-source models (GPT-4, Claude, Gemini) not evaluated
Reference
The takeaway: true code intelligence should be language-agnostic algorithmic reasoning, not memorization of one language's corpus. The data shows the industry is still far from that goal — and GPT-OSS-20B outperforming larger Qwen3 models cross-language suggests data distribution quality may matter more than parameter count.