How Independent Reviewers Judge Whether a Learning App Teaches Anything
App stores rank by downloads and star ratings, neither of which measures learning. The education reviewers who assess apps for schools use a different checklist — and it is worth knowing what is on it.

A five-star rating tells you that people who bothered to leave a review were pleased. A download count tells you a lot of people installed something once. Neither tells a parent or a teacher the only thing they actually want to know, which is whether a child who spends twenty minutes inside the app comes out knowing something they did not know before. Those two numbers are the ones every app store puts at the top of the page, and they are measures of popularity, not of learning.
There is a separate, much smaller world of reviewers who assess educational software the way a school would: against stated learning outcomes, with attention to how the app handles assessment, pacing and accessibility. Their checklists are public, and they are a genuinely useful thing for a parent to borrow — not because you need to run a formal evaluation on everything your child installs, but because knowing the five or six questions these reviewers ask will change what you look at in the first five minutes.
Does it name a learning outcome, or just a topic?
The first question is the one most apps fail. "Covers world history" is a topic. "Builds temporal awareness — the sense of what happened before what, and how far apart" is an outcome. The difference matters because an outcome can be checked and a topic cannot. Education reviewers typically work from a named list — academic relevance, reading, cognitive development, research skills, cultural awareness, self-direction, metacognition — and ask which of them the product actually serves, rather than which ones its marketing claims.
Metacognition is the one worth dwelling on, because it is the least visible and arguably the most valuable. It means the learner knowing what they do and do not know. A product that only ever tells a child they scored seven out of ten is not building it; a product that shows them which three they got wrong, and what the right answer was, is.
Is there assessment, and does it teach or just score?
This is the distinction between summative and formative assessment, and it is the one that separates a quiz that is decoration from a quiz that is part of the learning. A summative check produces a number at the end. A formative one feeds back into what the learner does next — it shows the specific gap, and it makes returning to the material the obvious move rather than a punishment.
There is a solid research reason to care. The testing effect, documented by Roediger and Karpicke in 2006 and replicated widely since, is the finding that being asked to retrieve information produces markedly better long-term retention than spending the same time re-reading it. Retrieval is not just a check on learning; it is one of the more effective forms of it. An app that ends a story with a few questions is, if the questions are any good, doing something more useful than one that simply offers the next story.
Can the learner set the pace, and change the conditions?
Self-paced learning and accessibility tend to be listed as separate criteria, but in practice they are the same question asked twice: can this be adapted to the person using it? Adjustable text size matters for a child with low vision and also for a tired adult reading at night. A choice of reading background — paper, cream, sand, dark — is a comfort setting for most people and a genuine accessibility requirement for some. Narration speed matters for a learner working in their second language.
Reviewers also look for something subtler: whether the interface language and the content language can be set separately. They are not the same preference. A Spanish-speaking parent may want to navigate in Spanish while their child reads in English, or the reverse, and a product that ties the two together has quietly decided that bilingual households do not exist.
Does it combine channels, or just stack them?
Multi-sensory learning is the most abused phrase on this list. Adding narration to text is not automatically a benefit — if the audio reads exactly what is on screen while a decorative image sits beside it, the learner is processing three versions of one thing and the redundancy can cost attention rather than add to it. What reviewers are actually looking for is whether the channels carry different parts of the meaning: the illustration showing what a place looked like, the narration carrying pace and emphasis, the text holding the detail you can go back to.
What this looked like when Wonder History was reviewed
The Educational App Store assessed Wonder History against this kind of framework and published its review in October 2026. It is a useful illustration precisely because it is not a press release — it lists what the reviewers thought worked, and a long section of what they thought should be fixed.
On the outcomes side they credited academic relevance, reading, cognitive development, research skills, cultural awareness, self-direction and metacognition. On features they named self-paced learning, multi-sensory learning, narration, formative assessment, accessibility and inclusion, and translation support. The three things they singled out as strongest were the breadth of the discovery system — browsing by timeline, geography, event type and historical figure rather than search alone — the narrated presentation, and the fact that the quiz shows a question-by-question breakdown rather than only a score.
Their summary sentence was that it is better characterised as a history exploration and knowledge-building platform than a complete history curriculum, which is a fair description and a useful one for a parent deciding what a thing is for. They were equally direct about the gaps: a catalogue large enough to feel dense for someone wanting a simple path, conventional multiple-choice questions, and age suitability that varies between individual stories and needs attention when choosing content. That last point is a real caution and worth repeating — a library spanning the whole of recorded history contains events that need an adult to decide when a particular child is ready for them.
The reviewers listed who they thought it suited: independent learners, students using it alongside formal study, parents choosing material for reading and listening at home, teachers using individual stories as introductions or discussion prompts, adult readers browsing by era and place, and multilingual households. Six quite different people, which is the practical consequence of separate Kids and Adults libraries sitting on the same catalogue.
The short version, for your next five minutes
Open the thing and ask: can I tell what it is trying to teach, rather than what it is about? When my child gets something wrong, does it show them what and why? Can I change the text size, the background, the narration speed, the language? Does the audio add something the text does not? And — the question no checklist includes but every parent should — would I be comfortable with them landing on any page of this at random? If the answer to that last one is no, the product needs a content filter you can trust, and you should go and find out whether it has one.
Frequently Asked Questions
What is the difference between formative and summative assessment?
A summative assessment measures what a learner knows at the end — a final score. A formative one feeds back into the learning while it is still happening: it identifies the specific gap and makes going back to the material the natural next step. A quiz that shows which questions you got wrong, and the correct answers, is formative. One that shows only "7/10" is not.
Is retrieval practice really better than re-reading?
For long-term retention, the research consistently says yes. Roediger and Karpicke's 2006 work found that learners who were tested on material retained substantially more of it later than those who spent the same amount of time re-reading, even though the re-readers felt more confident. Being asked to recall something is itself a form of learning, not just a measurement of it.
Why do reviewers care whether interface language and content language are separate?
Because they are different preferences and often belong to different people. In a bilingual household a parent may navigate comfortably in one language while the child reads in another. Tying the two settings together forces a single choice on a situation that genuinely has two.
Does adding narration to text always help?
No. If the audio reads exactly what is on screen and an image simply decorates it, the learner is processing the same information three ways, and the redundancy can compete for attention rather than reinforce. It helps when the channels carry different parts of the meaning — the picture showing the place, the voice carrying pace and emphasis, the text holding the detail.