A gamified mathematics module that pairs a large language model with exact graph-checking software has shown that AI-generated feedback can operate inside the rhythm of an ordinary middle school lesson, according to a new classroom study published in Heliyon. In two Korean schools, 49 students worked through a sequence of linear function tasks in which a friendly turtle character delivered staged hints whenever they went wrong, and the results offer one of the most detailed looks yet at how generative AI feedback actually behaves when real teenagers, real teachers, and a ticking 45-minute clock are all in the room.
The research, conducted by Sejun Oh, Haemee Rim, Seongkyeong Kim, and Yong-Oh Lee, tackles a stubborn problem in mathematics education. Linear functions demand that students coordinate equations, graphs, and changing quantities simultaneously, and decades of research show that learners struggle to translate between these representations and to explain why a line behaves as it does. Formative feedback, the kind that tells students not just whether they are right but what to do next, is exactly what this kind of reasoning needs. Yet providing it mid-lesson is enormously demanding for teachers, because every student stumbles at a different moment on a different idea.
The team’s answer was the AlgeoMath module, a system built on a deliberate division of labor. For tasks in which students constructed or manipulated graphs, correctness was determined by rule-based evaluators and the AlgeoMath application programming interface, which checked structured outputs such as slope, intercept, and relational constraints rather than interpreting screenshots. For written explanation tasks, a large language model, specifically gpt-3.5-turbo-0125 accessed through OpenAI’s Chat Completions API, generated staged feedback aligned with rubric criteria and moderated templates. No fine-tuning was performed; instead, the model was constrained through prompt templates, rubric criteria, staged output fields, and expert-reviewed feedback structures.
The gamification layer was deliberately restrained. A short narrative frame, progress indicators, and a turtle helper organized the sequence of 21 core tasks, and after an incorrect submission the turtle displayed a hint while students could retry on the same screen. Crucially, the designers avoided leaderboards, badges, rankings, and competitive scoring entirely. The game elements existed not to dangle rewards but to make the feedback cycle visible: students always knew where they were, what the turtle was suggesting, and what their next action could be. Hints came in three levels, from a light nudge to re-check part of an answer, through a stronger cue, to a final stage that could include a concise direct explanation enabling the student to proceed.
The classroom results were striking on the implementation side. All 22 ninth-graders at School A completed all 24 tasks, while the 27 eighth-graders at School B completed 98.1 percent of their 21-task sequence. Across both schools, roughly 70.7 percent of task instances were solved on the first submission, with a mean of about 1.53 submission events per task. More telling was what happened after failure: among initially incorrect answers that received a second attempt, 57.5 percent were corrected on that attempt, and among those still wrong after two tries, 43.3 percent were fixed on the third. The staged hint cycle, in other words, was not decoration; students were actually reading the feedback and using it to revise.
The module’s embedded scoring revealed a consistent pattern across both classrooms. Overall performance was high at School A, where the material served as enrichment after the main linear functions content had been taught, with a mean score of 96.1 out of 100, compared with 78.8 at School B, where eighth-graders met the material during the regular unit. But the rubric-dimension profiles told a more interesting story: in both schools, covariational reasoning and mathematical communication and justification were the weakest dimensions. Students could often manipulate graphs and identify slopes, but coordinating how two quantities change together, and explaining why equal slopes produce parallel lines, remained the hardest part. These profiles give teachers a concrete map of where follow-up instruction should go.
Students’ immediate self-reports also shifted. Of eight survey items administered before and after the lesson to 48 matched pairs, six showed statistically significant increases after Holm correction for multiple comparisons: enjoyment of learning mathematics, interest in mathematics, self-perceived competence, self-confidence, understanding difficult content, and perceived learning speed, with effect sizes ranging from 0.41 to 0.64. Notably, two general liking-and-interest items did not change significantly, a selective pattern the authors interpret as reflecting the lesson’s task experience rather than a wholesale transformation of attitudes toward mathematics. Written reflections echoed this, with students describing the game-like format as enjoyable and the turtle hints as helping them find their mistakes and keep going.
The study’s most sobering numbers concern the quality of the AI’s judgments. When 100 constructed responses were sampled and two experienced educators independently rated them as high, middle, or low against the same rubric, the educators agreed with each other 92 percent of the time, with a Cohen’s kappa of 0.863. The automated judgments agreed exactly with each educator only 69 percent of the time, with kappa values around 0.44 to 0.45. The disagreements concentrated in borderline responses with partially articulated reasoning, pinpointing a specific engineering target: the system needs clearer partial-credit guidance and exemplars of intermediate-quality explanations before its judgments can be trusted without human review.
A separate expert audit added a second layer of scrutiny. A panel of 20 mathematics educators conducted three rounds of design-time review of the feedback template library, which comprised 122 templates linked to the linear function tasks. Twenty-one templates, or 17.2 percent, were flagged at least once, and the dominant problem was not mathematical error, which accounted for only one flag, but scaffold calibration: 17 of the flagged issues involved messages whose level of help did not match their assigned stage, such as a Level 3 slot containing a reflective question rather than direct teaching. Encouragingly, only four flagged templates actually appeared in the pilot logs, accounting for just 2.1 percent of the 1,409 feedback instances delivered, and all templates and prompt rules were updated after the audit.
The authors are careful about the limits of what this pilot shows. The two classes were convenience-sampled, differed in grade level and task exposure, and there was no comparison condition, so no causal claims about learning gains can be made, and independent achievement and transfer were not measured. The logs show submissions after feedback appeared but cannot prove students attended to each hint. Still, the contribution is a concrete design template for teacher-supervised AI feedback: exact mathematical checks where determinism is possible, language-model hints where explanation matters, rubric-based reports for teachers, and human moderation as a standing safeguard. Future work, the team suggests, should prioritize partial-credit exemplars, scaffold calibration, and richer teacher analytics, alongside controlled comparisons with delayed outcome measures to test whether the immediate confidence boost translates into durable mathematical understanding.
Subject of Research: GPT-assisted formative feedback in a gamified linear function learning module for middle school mathematics
Article Title: GPT-assisted formative feedback for middle-school linear function inquiry: Design and classroom evaluation of a gamified AlgeoMath module
Article References: Oh, S., Rim, H., Kim, S., & Lee, Y.-O. (2026). GPT-assisted formative feedback for middle-school linear function inquiry: Design and classroom evaluation of a gamified AlgeoMath module. Heliyon, 12(15), Article e45558. https://doi.org/10.1016/j.heliyon.2026.e45558
Image Credits: AI Generated
DOI: 10.1016/j.heliyon.2026.e45558
Keywords: formative feedback, large language models, GPT, linear functions, gamification, mathematics education, middle school, covariational reasoning, rubric scoring, human-AI agreement, dynamic graphing, classroom study
Tags: classroom studycovariational reasoningdynamic graphingformative feedbackgamificationGPThuman-AI agreementLarge Language Modelslinear functionsMathematics educationmiddle schoolrubric scoring




