deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Khanmigo AI tutor trial in Tennessee schools finds modest math gains, weak engagement

A two-year randomized trial across 18 Tennessee middle schools found Khan Academy's Khanmigo AI tutor raised math scores by 0.06–0.08 standard deviations per year, but students rarely engaged it in substantive dialogue.

Khanmigo AI tutor trial in Tennessee schools finds modest math gains, weak engagement

A rare experimental test of AI tutoring

A two-year randomized controlled trial across 18 Tennessee middle schools has produced some of the first large-scale experimental evidence on whether generative AI can deliver on one of its most heavily promoted promises: a personal tutor for every student. According to the working paper, dated August 2026 and hosted on EdWorkingPapers, students who were randomly assigned to use Khan Academy alongside its AI tutor, Khanmigo, did improve their math achievement — but by amounts the researchers describe as comparable to what earlier studies found for Khan Academy practice without any AI component.

The study, which circulated widely after appearing on the Hacker News front page, used a cluster randomized design across the 18 schools. Assigned students used the tool during their existing daily remedial math sessions, with Khanmigo deliberately configured to coach students toward answers rather than supply them directly — the Socratic approach its makers advertise as the point of the product.

What the trial measured

The headline numbers are positive but modest. Assignment to the Khanmigo condition raised math achievement by 1.3 national percentile ranks per term, which the paper translates to roughly 0.06 to 0.08 standard deviations over a full school year. The researchers also estimate that a full year of active participation would imply an effect of about 0.14 standard deviations — a more meaningful gain, but one built on an assumption of sustained use that the study's own engagement data calls into question.

Crucially, the authors note that these gains resemble those previously documented for Khan Academy's ordinary practice exercises. In other words, the measurable benefit of adding an AI tutor on top of the platform appears, in this setting, to be close to what the platform already delivered without one.

Engagement emerged as the binding constraint

The study's most instructive findings concern how students actually behaved. Access was nearly universal: 96 percent of students tried Khanmigo at least once. But meaningful use was scarce. The median student sent messages to the tutor on only about a third of the days they spent practicing, and consulted it in only 17 percent of the exercise sessions in which they made a mistake — precisely the moments when a tutor should be most valuable.

When students did interact with Khanmigo, the exchanges were rarely the substantive mathematical dialogue the tool is designed to foster. According to the paper, most messages consisted of bare answers or clicks on suggested prompts rather than questions, reasoning or follow-up discussion. The technology was available and functioning; the students simply did not use it as intended.

The paper's conclusion is blunt on this point: the binding constraint appears to be engagement, not capability. Realizing the promise of AI tutoring, the authors argue, will require getting students to actually use the tutor — not just giving them access to it.

Why it matters

Most claims about AI tutoring rest on anecdotes, vendor case studies or short pilot programs. This is one of the first pieces of large-scale, multi-year randomized evidence, and it complicates the dominant narrative in both directions. The optimistic reading is that assignment caused real, measurable gains in math achievement at modest cost, and the implied effect of sustained participation — 0.14 standard deviations — would be substantial if it could be achieved. The sobering reading is that the AI component added little beyond what non-AI practice already produces, and that students largely declined to have the kind of conversations the product depends on.

For schools, districts and edtech developers, the practical lesson is that deployment is not adoption. A tool students open once and then ignore, or use by clicking suggested prompts, is functionally a different product from the one being sold. If AI tutoring is to move from hype to results, the hard problems may be behavioral and pedagogical — motivation, habits, and how tools prompt genuine mathematical thinking — rather than purely technical ones.

  • #ai-tutoring
  • #edtech
  • #khan-academy
  • #randomized-trial
  • #education

Related posts