"AI Slop" or "AI Enhancement"? 106 Hong Kong Students Answered with Real Grades
| Item | Detail | |------|--------| | Title | AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course | | Authors | David James Woo, Deliang Wang, Kai Guo | | arXiv | 2605.16275 (cs.CY, cs.AI, cs.CL, cs.MM) | | Date | April 2026, 23 pages | | Core contribution | First empirical study of whether AI-generated teaching materials are "slop" or "enhancement" — 106 Hong Kong EFL students; video + infographic preferences positively correlated with grades; high cognitive load negatively correlated with grades; weaker students spontaneously used AI materials as remedial scaffolds | | Link | https://arxiv.org/abs/2605.16275 |
First, the word "slop."
"AI slop" refers to the flood of low-quality, mass-produced AI content: slides spilling out like industrial wastewater, infographics generated like plastic flowers, podcasts looping like worn-out cassette tapes. The common trait: high volume, superficially plausible, needed by no one.
In education, "AI slop" carries a moral panic. Teachers worry students will use AI to write lazy assignments — but this post is about the reverse: what happens when teachers use AI to mass-produce learning materials? Can students tell "junk" from "good stuff"? Do the materials actually help learning?
The paper's title mirrors this question: AI Slop or AI-enhancement?
The answer: it depends on the design. And — the genuinely interesting part — not all AI materials are equal. Some types of students genuinely benefit; others are harmed by extra cognitive burden.
The Experiment: A Real Classroom of 106 Hong Kong EFL Students
This was a real academic English (EAP) course at a Hong Kong community college — not a lab experiment. The 106 learners of English as a foreign language faced what amounts to "a third task built on the difficulty of a second language."
The teacher used Google NotebookLM (a retrieval-augmented generation / RAG tool) to generate four types of AI materials from course content and student work:
- Videos: explanatory videos generated from course content
- Podcasts: audio explanations of course content
- Infographics: visual summaries
- Personalized feedback reports: individualized evaluations of each student's assignments
- Some students reported sharply higher intrinsic cognitive load with infographics — decoding the graphics themselves consumed attention
- Some reported extraneous cognitive load with podcasts — simultaneously imagining language, images, and holding academic terms in memory
The four media types were embedded in the normal course flow — not a special "AI experiment" class, but students' regular study plan. Afterwards, an explanatory sequential mixed-methods design was used: survey, then semi-structured interviews, then correlation analysis between preferences and course grades.
Key Findings: Video Preference Positively Correlates with Grades
Three layers of results:
1. Students generally found AI-generated materials useful and easy to use — standard Technology Acceptance Model dimensions. Average acceptance was high across all materials; nobody treated them as "slop."
2. Not all materials are equal. Students clearly preferred content tied to assessment (exam material turned into videos) and visual, multimodal formats. Videos and infographics were most popular. Podcasts got a lukewarm response — students said audio alone is hard to follow for academic content without visual support.
3. Video preference correlated positively with academic performance. Students who said "videos are most useful for me" scored significantly higher on final grades. It's a correlation, not causation — the paper did not experimentally show videos caused improvement — but the association was statistically significant.
Cognitive Load Is the Hidden Cost
Possibly the paper's most valuable finding:
High cognitive load was negatively correlated with grades.
If the media's design complexity is too high, processing cost eats the learning benefit. Although average acceptance was high, cognitive load distributions diverged:
Who Benefits Most?
The paper reports an important, under-developed finding:
Some lower-performing students spontaneously used AI materials as remedial scaffolds. They weren't told to do this — they reported it themselves in post-survey comments.
In interviews, these students described a pattern: opening videos after class and rewatching repeatedly — not one-time consumption but repeated use — treating AI explanations as a "pausable error-correction mechanism": pause, rewind, check the infographic, continue.
The authors word this cautiously — the sample is too small for statistical inference — but the observation is conceptually explosive: AI-enhanced materials don't uniformly raise everyone; they most improve the students who need them most. If replicated at scale, AI teaching materials could be an "anti-Matthew-effect" tool — the weak get more support, rather than the strong consolidating advantage.
Theoretical Limitations: What Can and Cannot Be Concluded
Several methodological constraints deserve candor:
1. The slop/enhancement binary is inadequate. Quality isn't a single dimension. An AI video may contain both serious biases (undetected cultural stereotypes) and highly accurate grammar explanations. The binary label obscures the hybrid nature of AI content.
2. Perception is not learning outcome. Conclusions rest mainly on student feelings ("I found this useful"). The only learning-outcome proxy is final grades — and the grade–preference correlation cannot be attributed to AI materials causing improvement.
3. The RAG baseline choice. Google NotebookLM was the only RAG tool. How much of "students liked AI materials" is due to NotebookLM's design, generation style, and output format rather than generic RAG features? A different tool could yield a very different acceptance profile, making conclusions non-robust to tool choice.
4. Sample homogeneity. 106 students at one Hong Kong community college, all EFL learners, cannot represent global education contexts. Trust in "authoritative" versus "AI-generated" materials may differ across educational cultures.
My Verdict
The paper's most valuable contribution is not the video–grade correlation — it is its rejection of the crude "AI slop" framework.
"AI slop" presumes AI-generated content is a quality floor that should be rejected on sight. The data show this may be entirely wrong in education. Students didn't treat AI materials as junk — they actively used them as learning scaffolds, and those who needed the most help benefited most.
The conclusion is not "AI-generated textbooks are good" — it's that their quality depends on whether the design accounts for cognitive layering, assessment alignment, and remediation. Drop AI into the classroom as a bulk content generator and you get slop. Design it intentionally as a personalized learning enhancer and you get enhancement.
Philosophically: perhaps "slop" and "enhancement" are not properties of the material — but of how it is inserted.
References
1. Woo, D.J., Wang, D., Guo, K. (2026). AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course. arXiv:2605.16275. 2. Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science. 3. Davis, F.D. (1989). Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology. MIS Quarterly. 4. Mayer, R.E. (2021). Multimedia Learning (3rd ed.). Cambridge University Press.