Research & Measurement · Northern Nigeria

What Happens When Teachers Grade Their Own Students' Reading?

📊  Skip to the interactive dashboard ↓

Teachers Don't Know What They Don't Know

How much should we trust a teacher's own sense of how well her students can read? Not a test score — just her gut feeling, built up over a term spent in the classroom with them.

Not much, it turns out. A 2024 study of thousands of teachers and students in India and Bangladesh found that teachers are strikingly bad at estimating their own pupils' skills. Researchers asked teachers to guess how their students would score on a test, then checked the guesses against the real results. The guesses barely tracked reality — nowhere close to how well teachers in wealthier countries size up their own students. And teachers weren't wrong at random. They oversold their weakest students the most. Nearly every teacher said they felt certain, or very confident, in guesses that turned out to be way off. Being wrong didn't feel wrong to them.

So What If You Take the Guessing Out of It?

That study was about open-ended judgment — a teacher forming an impression over a term, then putting a number on it. The obvious next question: what happens if you remove the guesswork? Hand teachers an actual standardized tool instead — a set passage, a stopwatch, a clear rubric — and ask them to run a real assessment, not just size a kid up.

A literacy programme across northern Nigeria gave us the answer. Teachers assessed pupils on a reading passage and sorted each one into a band based on how many words they read correctly per minute: couldn't read at all, just beginning, developing, or fluent. A few weeks later, independent supervisors quietly re-assessed the very same children, on the very same passage, using the very same four bands. If a standardized tool had fixed the problem, the two sets of scores should have matched closely.

They didn't.

The Tool Didn't Fix It

Teacher and supervisor scores lined up exactly a little more than four times out of ten. Strip out the agreement you'd expect from pure chance, and what's left barely counts as agreement at all — just one small step above a coin flip.

Look closer and the pattern is familiar. Teachers were good at spotting kids who genuinely couldn't read yet — they got that right about seven times out of ten. But for kids in the middle of the scale, starting to read but not yet fluent, teachers and supervisors barely agreed at all: about once in twenty-five times. The single most common mistake was calling a non-reader a "beginning reader" instead, which happened almost two out of every three times a teacher used that label. That's not random noise. It's a one-way drift — teachers consistently score their own pupils better than an outside observer does. Same direction as the earlier study. Except this time, teachers weren't guessing. They were following a script.

Why Does This Actually Matter?

These scores decide which kids get flagged for extra help. If a teacher's score makes a struggling reader look fine, that child slips through — and it's exactly the kids these programmes are built to catch who get missed.

There's no clean fix here. Teachers assessing their own pupils is fast, cheap, and keeps the people closest to the classroom in the loop. Outside supervisors are more accurate, but slower and too expensive to send in every week, everywhere. And because a standardized tool alone doesn't fix the bias, and because teachers who get this wrong tend to feel sure of themselves rather than unsure, you can't just ask a teacher to double-check her own work. Something outside the classroom has to catch the gap.

So What Happens Next?

That's part of why there's growing interest in tools that score reading fluency automatically — a recording of a child reading aloud, scored by software the same way every time, no matter who's holding the tablet. In theory, that's exactly the consistency missing here.

But it only works with teachers at the center of it. A machine can spit out a score. A teacher still has to turn that score into a plan — deciding what a struggling reader needs next and actually doing it in class. Push teachers to the sidelines and you've missed the point entirely. The goal was never a more accurate number for its own sake. It's giving teachers a number they can actually trust and act on.

Want to see the full breakdown — how the agreement was measured, exactly where teachers and supervisors disagreed? Keep reading for the full analysis, and explore the live dashboard below.

The Full Analysis

These findings come from comparing 1,844 pupils who were assessed by both their own teacher and an independent supervisor, on the identical reading passage and the same four fluency bands.

How agreement was measured

Agreement was measured with Cohen's kappa (κ), the standard statistic for comparing two raters scoring the same subjects into the same categories. It corrects raw percent-agreement for the agreement you'd expect from chance alone. Observed agreement here was 43.4%; expected agreement by chance was 36.6% — giving κ = 0.107, which matches the dashboard's own figure below. On the standard Landis & Koch scale, that falls into the "slight agreement" band.

Where teachers misassess most

The single largest error is teachers calling a pupil "Beginning" when the supervisor found the pupil could not read at all: 322 pupils, or 63.5% of every "Beginning" call a teacher made. Accuracy varies sharply by band:

Fluency bandTeacher accuracy
Cannot Read (most accurate)70.5%
Beginning20.1%
Developing (least accurate)4.1%
Fluent48.1%

A one-way drift, not random noise

Teachers don't miss in both directions equally — they consistently rate pupils higher than supervisors do. That's an optimistic skew, not a symmetric spread of mistakes, and it's the same direction found in the corroborating research below.

What we can't answer from this data

What percentage of teachers gave unreliable results? That can't be answered with a per-teacher number — the underlying data only has pooled and by-grade breakdowns, not a per-teacher or per-school one. Rather than invent a figure, this analysis uses the aggregate misassessment rate and directional bias above as the evidence for how widespread the problem is.

Corroborating research

This pattern lines up with Djaker, Ganimian & Sabarwal (2024), "Out of sight, out of mind? The gap between students' test performance and teachers' estimations in India and Bangladesh," Economics of Education Review, 102, 102575 (doi.org/10.1016/j.econedurev.2024.102575). That study found teacher estimates tracked real scores far more weakly in India and Bangladesh than in high-income countries, found teachers overestimated their weakest students by the widest margin, and found teachers were confidently wrong about it — the same shape of problem, in a different country, on a different task.

Explore the data yourself

Interactive Dashboard

Filter by grade, hover the flows, and dig into the full cross-tabulation and statistical tests below.

Assessment Flow
Teacher → Supervisor Fluency Category Flow
Flows show how pupils rated in each teacher category were rated by the supervisor. Hover any flow to highlight it.
Agreement
–
Observed
agreement

–
Expected
(chance)

–
Pupils
analysed
Cross-Tabulation
Teacher vs. Supervisor — Fluency Category Agreement
Rows = Teacher assessment  ·  Columns = Supervisor assessment  ·  Green cells = exact agreement (diagonal) · Cell % = row %
⚠  316 of 2,160 matched pupils excluded: teacher score recorded as "–" (not attempted or missing). Analysis n = 1,844.
Statistical Tests

Chi-Squared Test of Independence

Statistic (χ²)259.47
Degrees of freedom9
p-value< 0.001
N (all grades)1,844
The association between teacher and supervisor fluency ratings is highly significant — ratings are not independent. However, statistical significance alone does not indicate strong agreement; strength is measured by Kappa.

Cohen's Kappa — Inter-Rater Reliability

Observed agreement43.4%
Expected agreement (chance)36.6%
Kappa (κ)
0.107 Slight Agreement
Std. error0.014
Z-score7.44
p-value< 0.001
κ = 0.107 falls in the slight agreement range (κ 0.00–0.20, Landis & Koch 1977). Teacher and supervisor ratings agree only marginally above chance — indicating a meaningful reliability concern for this assessment instrument.