Summary
This paper by Liang Zhang, Yu Fu, and Xinyi Jin (arXiv:2603.25633, March 2026) investigates the relationship between large language models' (LLMs) mathematical problem-solving ability and their performance as step-level evaluators of student reasoning. As LLMs are increasingly deployed in mathematics education, they serve not only as problem solvers but also as graders of learners' reasoning processes. Using the GSM8K and MATH subsets of ProcessBench, the study compares evaluation accuracy on problems that models solve correctly versus incorrectly. The results show that evaluation accuracy is significantly higher on problems the models solve correctly than on those they solve incorrectly, with the difference statistically significant across two models and both datasets. This suggests a coupling between solving expertise and evaluation skill, implying that models' grading reliability in educational settings may depend on their own mastery of the underlying problems—an important consideration for deploying LLMs as automated reasoning evaluators.
Paper Overview
Field: Machine Learning
Authors: Liang Zhang, Yu Fu, Xinyi Jin
Published: 2026-03-26
arXiv: 2603.25633
Abstract
Large language models (LLMs) are increasingly used in mathematics education, not only as problem solvers but also as evaluators of learners' reasoning. This study uses the GSM8K and MATH subsets of ProcessBench to examine the relationship between mathematical problem-solving capability and step-level evaluation performance.
Key Findings
- On math problems that models solve correctly, evaluation accuracy is notably higher than on problems they solve incorrectly.
- The difference is statistically significant across two models and both datasets (GSM8K and MATH).
- This indicates that a model's evaluation reliability is coupled with its own problem-solving expertise on the given item.
Implications
- When LLMs act as graders of student reasoning, their judgments may be less trustworthy on problems outside their own competence.
- Deployments of LLM-based evaluation in education should account for the model's solving accuracy on the relevant problem distribution.
---
*Links: arXiv:2603.25633*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177169392