Twenty studies, and none in an American classroom

Lee V. Gaines, NPR

NPR visited an alternative school in Dayton, Ohio, where a small group of students spent a few weeks last year in morning conversations with a chatbot the district and its vendor built to sound like a discouraged teenager. The district paid SchoolAI $72,600 last year for several AI tools, the chatbot among them, and renewed for $78,650. The consultant hired to evaluate the pilot told NPR that every measure he tries stops at perception, meaning how students and staff say they feel about it.

The number under the whole piece comes from Stanford's SCALE Initiative. Its 2026 review went through more than 800 papers on AI in K-12 education and found twenty that can show whether a tool changed anything for students or teachers. None of the twenty was conducted with students in a U.S. K-12 classroom. Among the studies that do exist, the more promising results come from tools built with pedagogical guardrails, the kind that give hints and guide reasoning, rather than from general chatbots that hand over answers.

Wichita's AI specialist told NPR the district has not settled on what success would mean, so its student-facing pilots stay small. Wichita's caution is the defensible position in the piece. A district should agree on what it will measure before it buys the tool.

A school leader can ask a vendor which of the twenty studies its product resembles, and what it would take to find out whether the tool works with the district's own students. An experiment on students who are already struggling, as in Dayton, needs its measure settled before it begins, and the measure should be something other than how everyone feels afterward.

Our guide to questions to ask an AI vendor has the evidence questions in a form you can hand across the table.