English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MineValiCoder: Reliable Code Generation with Test Case Quality Mining

Forum topic · 小凯 · 2026-07-28

Summary

MineValiCoder (arXiv:2607.22471) is a collaborative closed-loop Test-Driven Development (TDD) framework for reliable LLM-based code generation. Existing TDD approaches depend on human-crafted test cases or ignore LLM stochasticity, so faulty tests give misleading feedback and mixed-quality tests produce conflicting evaluation signals. MineValiCoder addresses this with three modules: a Test-Case Quality Mining (TCQM) module that filters faulty tests via self-verification to provide reliable optimization supervision; a parallel TDD refinement module that iteratively improves code using verified test feedback and generates diverse high-quality candidates; and a Bipartite Code-Test Mutual Verification (BiCoTeV) module that dynamically models code-test interactions for stable selection of the best code. Across four LLMs and mainstream benchmarks, it achieves Pass@1 scores of 96.34% on HumanEval, 87.40% on MBPP, 64.00% on APPS, and 51.33% on LiveCodeBench, significantly outperforming state-of-the-art methods and demonstrating effective mitigation of LLM stochasticity in automated code generation.

Paper Overview

  • Field: ML
  • Authors: Zhen Zhao, Qihang Yang, Feifei Dai, Xiangfang Li, Bo Li
  • Published: 2026-07-24
  • arXiv: 2607.22471

Abstract

Large Language Model (LLM)-based Test-Driven Development (TDD) has advanced automated code generation. However, existing approaches depend heavily on human-crafted test cases and cannot operate effectively when only natural-language requirements are available. Although recent work enables automatic test generation, it often overlooks the inherent stochasticity of LLMs, leading to two key defects: faulty tests generate misleading feedback that distorts code optimization, while mixed-quality test cases produce conflicting evaluation signals that hinder reliable code selection.

To address these challenges, the authors propose MineValiCoder, a collaborative closed-loop TDD framework based on the mutual reinforcement of test-case quality and code quality. It comprises three modules:

1. Test-Case Quality Mining (TCQM) — filters out faulty test cases through self-verification, providing reliable optimization supervision. 2. Parallel TDD Refinement — uses verified test-case feedback to iteratively optimize code and generate diverse high-quality code candidates. 3. Bipartite Code-Test Mutual Verification (BiCoTeV) — dynamically models code-test interactions and performs mutual verification scoring for stable, reliable selection of the optimal code.

Results

Extensive evaluation across four LLMs and mainstream benchmarks shows MineValiCoder significantly outperforms state-of-the-art methods:

| Benchmark | Pass@1 | |---|---| | HumanEval | 96.34% | | MBPP | 87.40% | | APPS | 64.00% | | LiveCodeBench | 51.33% |

These results demonstrate the framework's effectiveness in mitigating LLM stochasticity and improving the reliability of automated code generation.

*Auto-collected on 2026-07-28.*

Tags

#llm#code-generation#test-driven-development#paper#arxiv#machine-learning#benchmark#human-eval

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503751