We share the interest others have in bridging the gap between the research community and instructional designers. Researchers are experts in measurement, and often have useful insights to provide, particularly to policymakers. But they rarely provide information on the "why," which allows a design team to make improvements to their program.
Most studies give an up or down measurement on a particular program, but rarely the underlying mechanisms that led to its "up-ness" or "down-ness." Instructional designers need the "why," and often rely on evidence that is perhaps less precise, or rigorous, than what the research community can provide. We think the answer is twofold: (i) try harder to wring utility out of higher-quality evidence, ((i) use that evidence as a conversation-starter, not a conversation-finisher.
An example: Lagos, Nigeria
How might one study, when combined with others, be used as a conversation starter? Let's look at Lipcan and Crawfurd's recent study comparing private schools, Bridge schools, and government schools in Lagos, Nigeria. We consider the results purely through the filter of program design — how the quality of teaching and learning may have impacted learning outcomes. This was not the focusing question of the research, which spans far wider: management practices, rate of corporal punishment, and return on investment calculations for Nigerian parents.
Three caveats
First, the test was administered one-third of the way through the school year. This is crucial because we can evaluate what had been taught, and what hadn't, by that time — helping us corroborate what we see in our own numeracy assessments and field officer observations.
Second, some of the skills assessed were not part of the 2nd grade national curriculum published by the Nigerian Educational Research and Development Council (NERDC), nor tied to the supplementary programs we have in place to remediate basic skill acquisition. The test was reasonably, but not purely, aligned to our program.
Third, the test was administered to second graders, and so has limited value for informing program design below or above that age group.
Two types of finding
The findings fall into two broad categories: skills where the program scores well, and skills where there are clear gaps. Where the program scored well, the data reinforces what we already see in our own assessments. Where there are gaps, the data prompts a productive question: is this a curriculum gap, a sequencing gap, or a time-on-task gap? Each diagnosis leads to a different design response.
This is the value of treating research evidence as a conversation-starter. Rather than reading the study as a verdict on the program, we read it as a prompt. It surfaces questions that our own internal data can partially answer — and points us toward the specific design changes most likely to move the needle.
The research community can play a vital role not just in measuring impact, but in helping design teams understand the mechanisms behind their results. We hope this example encourages more researchers to present findings with that practical utility in mind — and more practitioners to read research with that kind of interrogative curiosity.