A PNAS Study with Turkish Pupils Reveals AI's Real Impact in the Classroom
Have you ever sensed a disconnect between how your pupils perform on their homework and their exam results? Everything seems to go smoothly throughout the lesson: questions get solved, homework gets handed in. Then the exam arrives, and you find the picture looks rather different…
A study conducted in recent years with secondary school pupils in Türkiye, published in PNAS, one of the world’s most prestigious scientific journals, has put numbers on exactly why this disconnect happens.
Unexpected Results from the Experiment
The researchers divided roughly a thousand Turkish secondary school pupils into three groups for a mathematics study. The first group used a standard AI, that is, a system where they could type in the question and get the answer directly if they wished. The second group used a pedagogically designed AI that did not give the answer outright; this system offered hints, asked questions and waited for the pupil to work out the next step themselves. The third group used no AI at all and worked only with the textbook.
Interestingly, during the practice phase both AI groups pulled ahead of the textbook group. Pupils using the standard AI improved their practice performance by 48 per cent. Those using the hint-giving AI improved by 127 per cent. There was a substantial gap between the two groups in practice, and the hint-giving AI was leading by a clear margin!
Then came the exam stage. With all the learning done, an environment was created that was completely isolated from AI, where only the pupil and the question paper remained, and pupils from all three groups sat the same exam. The results were quite startling: the standard AI group performed 17 per cent worse on the exam than the group that had worked with the textbook, while the hint-giving AI group scored the same as the textbook group, which meant the gains made in practice had been successfully carried through to the exam, thanks to the right kind of AI use!

The crutch effect
The researchers call this the “crutch effect”. The pupil types their question into the standard AI, and the system lays out the solution step by step. The pupil reads it, understands it, perhaps even takes notes. But what is really happening in that process? The pupil never personally experiences the path of thinking inside that solution: where to turn at which point, where they might get stuck, and how to get out of that sticking point…
The study brings a striking detail to light here as well. Of the pupils using the standard AI, 67 per cent either copied and pasted the question directly into their conversations or simply said “give me the answer”. In the hint-giving AI group, that figure stayed at 37 per cent. What is more, the researchers also demonstrate that a large proportion of pupils did not even read the solution the AI provided; they copied it straight over.
What stands out here is this: both systems in the experiment used the same AI model, namely GPT-4. The only difference between them was the design. The hint-giving system contained error patterns and correct solutions prepared by teachers. It knew where pupils tended to go wrong and how, and it guided them accordingly. In other words, the difference lay not in the technology but in the pedagogical knowledge behind the technology!
With the hint-giving AI, something different happens at precisely this point. Instead of giving the answer, the system nudges the pupil to think, points out a direction, asks a question and then waits. The pupil has to grapple with the problem, sometimes getting stuck, sometimes making mistakes. It sounds slow, but learning science says quite the opposite and calls this “desirable difficulties”. Feeling that something is easy does not mean learning has taken place. Real learning sometimes happens right in the middle of that struggle.
The Key Takeaway: AI by Itself Is Not Enough for Education
In many schools today, permission to use AI still amounts to a single sentence, yet none of the questions of which AI, for what purpose and how ever get answered. And the enormous gap this research reveals is born precisely in that void.
That Is Why Madlen’s Difference Lies Not in the Technology, but in the Design
What this research tells us is this: what an AI says to a pupil matters far more than what it thinks on their behalf. We built Madlen on exactly this understanding. In Madlen Okul, when pupils ask a question, the system never gives them the answer; it asks, it guides, and it waits for the pupil to work out the next step themselves. This Socratic approach we have embraced carries the essence of a pedagogical tradition thousands of years old: real learning comes not from receiving the right answer, but from confronting the right question!

To understand why this matters, one only needs to look at learning science. When a pupil truly wrestles with a concept, wanders down the wrong paths and finds their own way out, that knowledge leaves a much deeper imprint on their mind.
However powerful AI may be, knowledge that is not constructed in the pupil’s own mind does not stay there. Technology does not change this truth; it only makes it more visible. In Madlen, every hint, every question, every moment of waiting exists for this reason. Because we know that making room for pupils to think is always more valuable than thinking on their behalf!

Sources:
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122