Teachers Don't Know What They Don't Know
How much should we trust a teacher's own sense of how well her students can read? Not a test score — just her gut feeling, built up over a term spent in the classroom with them.
Not much, it turns out. A 2024 study of thousands of teachers and students in India and Bangladesh found that teachers are strikingly bad at estimating their own pupils' skills. Researchers asked teachers to guess how their students would score on a test, then checked the guesses against the real results. The guesses barely tracked reality — nowhere close to how well teachers in wealthier countries size up their own students. And teachers weren't wrong at random. They oversold their weakest students the most. Nearly every teacher said they felt certain, or very confident, in guesses that turned out to be way off. Being wrong didn't feel wrong to them.
So What If You Take the Guessing Out of It?
That study was about open-ended judgment — a teacher forming an impression over a term, then putting a number on it. The obvious next question: what happens if you remove the guesswork? Hand teachers an actual standardized tool instead — a set passage, a stopwatch, a clear rubric — and ask them to run a real assessment, not just size a kid up.
A literacy programme across northern Nigeria gave us the answer. Teachers assessed pupils on a reading passage and sorted each one into a band based on how many words they read correctly per minute: couldn't read at all, just beginning, developing, or fluent. A few weeks later, independent supervisors quietly re-assessed the very same children, on the very same passage, using the very same four bands. If a standardized tool had fixed the problem, the two sets of scores should have matched closely.
They didn't.
The Tool Didn't Fix It
Teacher and supervisor scores lined up exactly a little more than four times out of ten. Strip out the agreement you'd expect from pure chance, and what's left barely counts as agreement at all — just one small step above a coin flip.
Look closer and the pattern is familiar. Teachers were good at spotting kids who genuinely couldn't read yet — they got that right about seven times out of ten. But for kids in the middle of the scale, starting to read but not yet fluent, teachers and supervisors barely agreed at all: about once in twenty-five times. The single most common mistake was calling a non-reader a "beginning reader" instead, which happened almost two out of every three times a teacher used that label. That's not random noise. It's a one-way drift — teachers consistently score their own pupils better than an outside observer does. Same direction as the earlier study. Except this time, teachers weren't guessing. They were following a script.
Why Does This Actually Matter?
These scores decide which kids get flagged for extra help. If a teacher's score makes a struggling reader look fine, that child slips through — and it's exactly the kids these programmes are built to catch who get missed.
There's no clean fix here. Teachers assessing their own pupils is fast, cheap, and keeps the people closest to the classroom in the loop. Outside supervisors are more accurate, but slower and too expensive to send in every week, everywhere. And because a standardized tool alone doesn't fix the bias, and because teachers who get this wrong tend to feel sure of themselves rather than unsure, you can't just ask a teacher to double-check her own work. Something outside the classroom has to catch the gap.
So What Happens Next?
That's part of why there's growing interest in tools that score reading fluency automatically — a recording of a child reading aloud, scored by software the same way every time, no matter who's holding the tablet. In theory, that's exactly the consistency missing here.
But it only works with teachers at the center of it. A machine can spit out a score. A teacher still has to turn that score into a plan — deciding what a struggling reader needs next and actually doing it in class. Push teachers to the sidelines and you've missed the point entirely. The goal was never a more accurate number for its own sake. It's giving teachers a number they can actually trust and act on.
Want to see the full breakdown — how the agreement was measured, exactly where teachers and supervisors disagreed? Keep reading for the full analysis, and explore the live dashboard below.