Paper Overview
- Field: ML
- Authors: Zhen Zhao, Qihang Yang, Feifei Dai, Xiangfang Li, Bo Li
- Published: 2026-07-24
- arXiv: 2607.22471
Abstract
Large Language Model (LLM)-based Test-Driven Development (TDD) has advanced automated code generation. However, existing approaches depend heavily on human-crafted test cases and cannot operate effectively when only natural-language requirements are available. Although recent work enables automatic test generation, it often overlooks the inherent stochasticity of LLMs, leading to two key defects: faulty tests generate misleading feedback that distorts code optimization, while mixed-quality test cases produce conflicting evaluation signals that hinder reliable code selection.
To address these challenges, the authors propose MineValiCoder, a collaborative closed-loop TDD framework based on the mutual reinforcement of test-case quality and code quality. It comprises three modules:
1. Test-Case Quality Mining (TCQM) — filters out faulty test cases through self-verification, providing reliable optimization supervision. 2. Parallel TDD Refinement — uses verified test-case feedback to iteratively optimize code and generate diverse high-quality code candidates. 3. Bipartite Code-Test Mutual Verification (BiCoTeV) — dynamically models code-test interactions and performs mutual verification scoring for stable, reliable selection of the optimal code.
Results
Extensive evaluation across four LLMs and mainstream benchmarks shows MineValiCoder significantly outperforms state-of-the-art methods:
| Benchmark | Pass@1 | |---|---| | HumanEval | 96.34% | | MBPP | 87.40% | | APPS | 64.00% | | LiveCodeBench | 51.33% |
These results demonstrate the framework's effectiveness in mitigating LLM stochasticity and improving the reliability of automated code generation.
*Auto-collected on 2026-07-28.*