Do reading apps actually work?
Yes, a little. Less than the people selling them suggest, and well short of a teacher. Here is the evidence with the unflattering parts left in.
The short answer
Reading software helps. The effect has been measured across more than a hundred studies and it is real. It is also modest, and it shrinks by roughly a third to a half once you count only the tests that were not written by the people running the study.[1]
No reading app on the market has evidence that it can teach a child to read on its own. Not ours either. Anyone who tells you otherwise is describing a product, not a result.
The biggest study, and what it found
In 2025 a team at Stanford and the University of Amsterdam pooled 119 studies of literacy software, drawn from 105 published papers, covering children from kindergarten through grade 5 in work published between 2010 and 2023.[1] It is the broadest look anyone has published at this question.
Their results come as effect sizes. An effect size puts findings from different studies on one scale, so you can compare them: a small number means a small real difference, a bigger number means a bigger one. Across every test used in those studies, the pooled effects were 0.33 for decoding, 0.30 for language comprehension, 0.23 for reading comprehension and 0.81 for writing.[1]
The authors call these medium to large by the standards of education research, and note that they compare well with education interventions generally.[1] That is a genuinely good finding and we are not going to talk it down.
Then look at who wrote the test
Many studies of educational software measure the result with a test the research team built themselves. Some of those tests ask about exactly the words and letter sounds the program had just taught. The reviewers call these proximal measures, and they warn readers off them, quoting the What Works Clearinghouse handbook published by the US Department of Education: “When outcome measures are closely aligned with or tailored to the intervention condition, the study findings may not be an accurate indication of the effect of the intervention.”[1][3]
So they ran the numbers again using standardized tests only, the established ones with published evidence of reliability across many studies. Everything got smaller.[1]
| Skill | All tests | Standardized only |
|---|---|---|
| Decodingsounding words out | 0.33 | 0.23 |
| Language comprehensionunderstanding spoken language | 0.30 | 0.12 |
| Reading comprehensionunderstanding what you have read | 0.23 | 0.14 |
| Writingonly 6 studies measured this | 0.81 | 0.34 |
Language comprehension lost about sixty percent of its effect. Reading comprehension, the thing everyone actually cares about, landed at 0.14.[1] These are the numbers you will not find on a vendor's landing page, and they are the ones worth knowing.
Set against a person teaching
The reference figure for systematic phonics taught by a teacher is 0.41, from the meta-analysis behind the US National Reading Panel: 66 treatment and control comparisons drawn from 38 experiments.[2] It was larger when phonics started early, 0.55 in kindergarten and first grade, and smaller when it started later, 0.27.[2]
Put 0.41 next to 0.23 on a standardized decoding test and the ordering is hard to miss. Be careful with that arithmetic anyway. The reviewers say so themselves: “Differences across meta-analyses make comparisons fraught.”[1] Different studies, different decades, different comparison groups. Treat it as a rough ordering rather than a scoreboard.
One more figure from the same review, and it is the sharpest. It points to a separate synthesis of programs for struggling elementary readers in which technology-supported adaptive instruction, which is the category most reading apps fall into, came out at 0.09 and was not statistically significant.[1][4]
Four things nobody has measured
A meta-analysis can only pool the studies that exist. Here is what is missing from the pile.
- Almost all of it is school software. To be included, an intervention had to be something that could run in a school setting.[1] A child on a sofa at home is a different situation, and it has not been studied at anything like this scale.
- Almost nothing was followed up. 19 of the 119 studies checked whether the gains were still there later.[1] Effects that fade are counted the same as effects that last.
- Almost nothing measured how children felt. 10 of the 119 studies included any measure of motivation or engagement at all.[1]
- Nobody has tested us. Plimmery is built on the phonics evidence above, but that is a claim about our method, not a measured result for our app. If a trial is ever run, we will publish it here whichever way it comes out.
So what is an app for?
The repetitive part. A child needs hundreds of short goes at letter sounds, each one answered instantly and patiently, and that is a thing software is genuinely good at and most adults are not. What software cannot do is notice that your child has gone quiet, or read them a story, or care.
A supplement, then. Never a replacement for teaching, and never a replacement for you.
References
Every source here was opened and read. Where we could only reach a finding second-hand, the entry says so.
- Silverman, R. D., Keane, K., Darling-Hammond, E., & Khanna, S. (2025). The Effects of Educational Technology Interventions on Literacy in Elementary School: A Meta-Analysis. Review of Educational Research, 95(5), 972-1012. DOI: 10.3102/00346543241261073 Open sourceRead in full, from the open-access copy in the University of Amsterdam repository.
- Ehri, L. C., Nunes, S. R., Stahl, S. A., & Willows, D. M. (2001). Systematic Phonics Instruction Helps Students Learn to Read: Evidence from the National Reading Panel's Meta-Analysis. Review of Educational Research, 71(3), 393-447. DOI: 10.3102/00346543071003393 Open source
- What Works Clearinghouse (2022). Procedures and Standards Handbook, Version 5.0. Institute of Education Sciences, U.S. Department of Education. Open sourceThe sentence we quote was read where reference 1 quotes it, on page 27 of the handbook.
- Neitzel, A. J., Lake, C., Pellegrini, M., & Slavin, R. E. (2022). A Synthesis of Quantitative Research on Programs for Struggling Readers in Elementary Schools. Reading Research Quarterly, 57(1), 149-179. DOI: 10.1002/rrq.379 Open sourceThe full text is behind a paywall we could not open. The 0.09 figure is quoted as reference 1 reports it, not from the original.