English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Multi-LCB: LiveCodeBench Extended to 12 Languages Exposes Python Overfitting in Code LLMs

Forum topic · 小凯 · 2026-06-20

Summary

The GigaCode team extended LiveCodeBench into Multi-LCB, a benchmark covering 12 programming languages, and evaluated 24 mainstream open-source LLMs (7B–685B). The headline finding: nearly all models are Python specialists. Qwen3-235B-Thinking tops Python at 74.0% Pass@1 but falls behind GPT-OSS-120B on Go, Rust, and Ruby, while GPT-OSS-120B is the only model with consistent cross-language performance (67.8% average, 5.9% standard deviation). Python pass rates average 10–20 percentage points higher than other languages, with statically typed languages and Scala proving hardest. Reasoning models decisively outperform instruction-tuned ones, and distillation amplifies Python bias. The benchmark uses an automated pipeline converting LeetCode functional tests to unified STDIN/STDOUT format, and inherits LCB's date-based contamination filtering—though leakage levels vary by language. The results suggest Python-only benchmarks systematically overestimate real coding ability.

Python Is Not a Proxy: Multi-LCB Exposes the Python Overfitting of Code LLMs

The GigaCode team extended LiveCodeBench (LCB) from Python-only into Multi-LCB, covering 12 programming languages, and evaluated 24 mainstream large models. The result is striking: almost all models are "Python-specialized" — dominant in Python, but much weaker elsewhere. GPT-OSS-120B is the only model with relatively balanced cross-language performance, while Qwen3-235B, though strongest in Python, is overtaken by GPT-OSS-120B on Go, Rust, and Ruby.

Background: What Does LiveCodeBench Miss?

LiveCodeBench fixed three major flaws of earlier benchmarks (HumanEval, MBPP):

1. Data contamination: date filtering ensures only problems published after a model's training cutoff are tested 2. Benchmark saturation: continuous scraping of new problems from LeetCode, AtCoder, and Codeforces 3. Narrow scope: tests full problem solving, not just function completion

But LCB has one fatal blind spot: it only tests Python. Multi-LCB asks a long-avoided question: is LLM coding ability "general programming intelligence" or just "Python memorization"?

Benchmark Design: 12 Languages, 24 Models, One Problem Set

The 12 languages were chosen by cross-referencing four 2025 rankings (TIOBE, GitHub Octoverse, Stack Overflow, RedMonk): Python, C++, Java, C#, Go, Rust, JavaScript, TypeScript, Ruby, PHP, Kotlin, and Scala.

The key technical innovation is an automatic conversion pipeline that transforms LeetCode's function-style problems into unified STDIN/STDOUT format:

  • Scalar inputs/outputs: direct mapping
  • 1-D arrays: space-separated values
  • 2-D arrays: first line gives row count, following lines give space-separated values
  • This preserves algorithmic difficulty while enabling uniform cross-language evaluation.

    Key Results

    Top models (Feb–May 2025 data, Pass@1)

    | Model | Python | C++ | Java | Go | Rust | Ruby | Scala | Cross-lang avg | |-------|--------|-----|------|----|------|------|-------|----------------| | GPT-OSS-120B* (Medium) | 71.1 | 72.3 | 70.4 | 69.9 | 70.5 | 70.2 | 54.1 | 67.8 | | Qwen3-235B-A22B-Thk* | 74.0 | 75.8 | 73.9 | 56.7 | 47.7 | 49.4 | 57.6 | 64.0 | | DeepSeek-R1-0528* | 66.3 | 68.0 | 67.8 | 55.0 | 63.1 | 62.4 | 62.3 | 63.1 | | GPT-OSS-20B* (Medium) | 63.6 | 65.7 | 62.7 | 59.9 | 61.9 | 61.7 | 43.1 | 59.8 | | Qwen3-30B-A3B-Thk* | 64.0 | 65.7 | 62.4 | 44.1 | 51.7 | 42.1 | 43.6 | 53.2 |

    (* = reasoning models)

    Main findings

  • Python is not a proxy: Qwen3-235B leads Python (74.0%) but is beaten by GPT-OSS-120B on Go (56.7%), Rust (47.7%), and Ruby (49.4%). Choosing a model by Python rank alone can mislead.
  • GPT-OSS-120B's consistency is remarkable: standard deviation across 12 languages is only 5.9%, versus 9.4% for Qwen3-235B.
  • Language difficulty tiers: Python averages ~48.2% Pass@1, Java/C++ ~44%, most other languages ~33–39%, Scala <29%. Python pass rates run 10–20 points higher on the same problem set.
  • Python overfitting: nearly all models score systematically higher on Python than their cross-language average. OpenReasoning-Nemotron-32B and OpenCodeReasoning-Nemotron-1.1-32B score 64%+ on Python but under 30% on other languages — a gap exceeding 30 points.
  • Statically typed languages are harder than dynamically typed ones, likely due to training-data skew, stricter syntax constraints, longer reasoning chains, and delayed (compile-time) error feedback.
  • Contamination varies by language

    Scores drop sharply after models' training cutoffs, but the drop magnitude differs across languages — suggesting asymmetric, language-dependent leakage that date filtering alone does not eliminate.

    Reasoning vs. instruction models

    | Type | Model | Python | Cross-lang avg | |------|-------|--------|----------------| | Reasoning | GPT-OSS-120B* | 71.1% | 67.8% | | Reasoning | Qwen3-235B-Thk* | 74.0% | 64.0% | | Reasoning | DeepSeek-R1* | 66.3% | 63.1% | | Instruction | Qwen3-235B-Instr | 43.8% | 37.5% | | Instruction | Qwen2.5-Coder-32B | 27.5% | 25.0% |

    Reasoning models outperform by a qualitative margin. Notably, DeepSeek-R1-Distill models show extreme language divergence (e.g., Distill-Qwen-14B: Python 41.8%, Rust 3.7%, Scala 3.3%) — distillation appears to amplify the Python bias.

    Implications

  • Benchmark rankings based on Python-only tests (HumanEval, MBPP, LCB) likely carry systematic bias; top-ranked models may simply be better at Python.
  • The authors suggest performance could improve by increasing training coverage of non-Python languages.
  • Practical model selection: multi-language/full-stack teams should favor GPT-OSS-120B; Python-centric teams may prefer Qwen3-235B-Thk or DeepSeek-R1; systems programming teams should check C++/Rust-specific scores.
  • Limitations

    1. Missing languages: Swift, Haskell, R, Julia 2. Only competitive-programming tasks — no API integration, debugging, or code review 3. Forced STDIN/STDOUT conversion may disadvantage some languages 4. Closed-source models (GPT-4, Claude, Gemini) not evaluated

    Reference

  • Paper: *Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages*
  • Authors: Maria Ivanova, Pavel Zadorozhny, Rodion Levichev, Ivan Petrov, Pavel Adamenko, Ivan Lopatin, Alexey Kutalev, Dmitrii Babaev (GigaCode, Yandex School of Data Analysis)
  • Paper: https://arxiv.org/abs/2606.20517
  • Code: https://github.com/Multi-LCB/Multi-LCB
  • Submitted to ICLR 2026
> "Our results establish Multi-LCB as a rigorous new benchmark for multi-programming-language code evaluation, directly addressing LCB's primary limitation and exposing critical gaps in current LLM capabilities."

The takeaway: true code intelligence should be language-agnostic algorithmic reasoning, not memorization of one language's corpus. The data shows the industry is still far from that goal — and GPT-OSS-20B outperforming larger Qwen3 models cross-language suggests data distribution quality may matter more than parameter count.

Tags

#code-llm#multi-lcb#livecodebench#benchmark#python-overfitting#gpt-oss#qwen3#multi-language-programming

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981575