English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Million Tutoring Moves (MTM) Dataset Opens Up 4,654 Real Math Tutoring Transcripts for AI Tutoring Research

Forum topic · 小凯 · 2026-05-18

Summary

Researchers from Cornell, Stanford, MIT, and CMU — including Kizilcec, Vanacore, Zhou, Justin Reich, and Ken Koedinger — have released the first version of the Million Tutoring Moves (MTM) dataset, addressing a long-standing bottleneck in AI tutoring research: the scarcity of high-quality, real tutoring dialogue data. The initial release contains 4,654 complete one-on-one math tutoring session transcripts from a US non-profit online tutoring platform, capturing real interactions such as questioning, explaining, hinting, correcting, and encouragement — not lab recordings. The project's long-term goal is one million tutoring moves, guided by design principles of open licensing, safety via de-identification and privacy review, large scale, broad coverage across grade levels and tutoring strategies, and multimodality (text plus behavioral logs, with possible speech and canvas data). Open questions remain: current coverage is math-only, tutoring effectiveness labels are unclear, and residual identifiability risks after de-identification are not fully specified. Reference: arXiv:2605.08092 [cs.CY].

AI intelligent tutoring system research faces an awkward bottleneck: high-quality, real human–human or human–AI tutoring dialogue data is hard to obtain. Public datasets are either too small (a few hundred dialogues), lack multimodal information (text only, no speech or whiteboard behavior), or cannot be released for privacy reasons. As a result, most AI tutoring systems are effectively trained on synthetic data.

Kizilcec, Vanacore, Zhou, and collaborators from Cornell, Stanford, MIT, and CMU (including Justin Reich and Ken Koedinger) have released the first version of the Million Tutoring Moves (MTM) dataset — 4,654 math tutoring dialogue transcripts from a non-profit online tutoring platform in the United States. Each transcript contains the complete interaction between tutor and student: questioning, explaining, hinting, correcting, encouraging. This is not lab-recorded data — it is real one-on-one online tutoring.

Key points

  • Scale: 4,654 real tutoring session transcripts in version 1; the long-term goal is "million-level" tutoring moves — far beyond the current release.
  • Authenticity: genuine one-on-one online tutoring sessions, not staged laboratory recordings.
  • Design principles highlighted in the paper:
  • Openness: open-source license without usage restrictions
  • Safety: de-identification and privacy review
  • Scale and breadth: coverage across grade levels, subjects, and tutoring strategies
  • Multimodality: text + behavioral logs + possible speech and canvas data
  • Open questions

  • Domain coverage: the 4,654 dialogues currently cover mathematics only — how broad is the eventual coverage?
  • Effectiveness labels: tutor quality may vary considerably; does the dataset annotate the outcome or effectiveness of each session?
  • De-identification depth: after student names and school information are removed, do indirectly identifying details remain in the dialogue text?

References

1. Kizilcec, R., Vanacore, K., Zhou, Z., et al. (2026). *Million Tutoring Moves (MTM): An Open Multimodal Dataset for the Science of Tutoring*. arXiv:2605.08092 [cs.CY]. 2. Chi, M. T. H., et al. (2001). *Learning from Human Tutoring*. Cognitive Science. 3. Koedinger, K. R., et al. (2012). *Data Mining and Education*. Wiley Interdisciplinary Reviews: Cognitive Science.

Tags

#ai-tutoring#dataset#education-technology#machine-learning#multimodal-data#research#privacy#open-data

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620322