
Personalised learning at class scale: what AI changes
Sit through enough education technology pitches and you will hear about Bloom's two sigma. The claim is that a child tutored one to one performs two standard deviations better than a child in an ordinary classroom, which would put an average student above the 98th percentile. It is the intellectual foundation of roughly every "AI tutor for every child" deck of the last three years.
It is worth knowing what the study actually was, because the details change what personalised learning can realistically be built to do.
What the two sigma study actually was
Benjamin Bloom published the paper in Educational Researcher in 1984, drawing on dissertation work by two of his doctoral students. The specifics, set out at length by No More Marking and summarised in the Wikipedia account:
- The intervention ran for about three weeks, eleven lessons of forty minutes, with testing immediately afterwards.
- The subjects were cartography and probability, chosen partly because they were unfamiliar. Beginners improve fast on new material, then plateau.
- The tests were designed by the researchers rather than standardised.
- A couple of hundred students took part.
- It has never been replicated.
And the detail that matters most: the tutored group did not only get tutoring. They also received extra corrective feedback and corrective tests that the comparison group did not. Three variables moved at once, and the headline attributes the whole effect to one of them.
Post-Covid, several governments funded large catch-up tutoring programmes with this paper somewhere in the justification. The results came back nothing like two sigma. That is not surprising given what the paper actually measured, but it did cost a lot of money to discover.
None of which makes tutoring bad. It makes two sigma from one-to-one contact an unsupported number, and it means the interesting question was never "how do we give every child a tutor".
The part everybody skips is the useful part
Look again at what the tutored condition contained. Mastery learning, corrective feedback, and corrective tests. Then ask which of those actually require a one-to-one setting.
None of them do.
Mastery learning is a sequencing decision: you do not move on until the thing is secure. Corrective feedback is a marking decision: the response tells the learner what specifically to fix. Corrective testing is a checking decision: you come back and verify the gap closed. All three are properties of the loop, not of the staffing ratio.
What one-to-one really buys is proximity. A tutor sitting beside one child can see, in real time, what that child cannot do yet, and act on it immediately. The ratio is a delivery mechanism for a diagnosis.
Which reframes the problem for a class of forty. You are not trying to buy forty tutors. You are trying to get the diagnosis without the tutor.
Why this has been out of reach, honestly stated
Teachers have always known which children need what. The constraint has never been insight, it has been arithmetic.
To personalise properly you need three things in sequence: a per-skill picture of every child, material aimed at each gap, and a check that the gap closed. For thirty children across a term, the first is a marking workload nobody has, the second is a preparation workload nobody has, and the third almost never happens at all. Most differentiation in practice collapses into three worksheets labelled by ability, set once, never re-checked.
What AI changes here is the arithmetic. A loop that good teachers have always known how to run becomes survivable across thirty children, which is a smaller claim than teaching better than a teacher and a more useful one.
Let the marking decide who gets what
Teaching each child like the only one in the room is the right goal. The useful question is what has to happen for it to be real by Tuesday morning rather than aspirational.
It starts with not deciding in advance. A gap sometimes belongs to one child and nobody else, and that child gets a piece written for them alone. More often four or six turn out to be missing the same thing, and they get the same practice because they genuinely need the same thing. Who receives what follows from the diagnosis, rather than from a grouping fixed at the start of term.
What it never needs to be is thirty bespoke worksheets produced for the sake of it. That is easy to generate now and it is still the wrong shape: an enormous amount of material for one teacher to check, and unchecked AI-generated work going straight to children is not something we are willing to ship.
So Zippy drafts practice for a skill and whoever the marking says is carrying that gap, whether that is one student or seven. From where a child sits, nothing is diluted either way: the work in front of them is aimed at what their own marking found. What changes on your side of the desk is that you are reviewing what the class actually needs rather than authoring thirty things from nothing.
That is what makes catering to every student survivable. Not a tutor each, and not one worksheet for everyone, but each child getting the next thing they specifically need, from a teacher who had the time to check it.
The loop, concretely
What this looks like in the product, with no steps hidden:
- Mark the set you were marking anyway. Every piece is scored per criterion and per skill, with comments tied to the sentence that earned them.
- The skills accumulate into a record. Each child gets a skill-by-skill picture that builds over the term: secure here, developing there, missed in three of the last four pieces.
- Zippy drafts practice aimed at what the marking found, for one student or for everyone carrying that gap.
- You read it and change it. Nothing is assigned without you. This is the step that cannot be automated away and we are not trying to.
- It comes back marked against the same skills, so you can see whether the gap actually closed.
Step five is the corrective testing from the 1984 study, and it is the step almost every "personalised learning" product quietly omits. Generating differentiated worksheets is easy. Closing the loop on whether they worked is the part that makes it learning rather than activity.
Students see their half of it too: their own skills, their feedback, and the task set next. A mark tells a child how they did. A skill breakdown tells them what to do about it, which is a more useful thing for an eleven-year-old to be holding.
What we do not claim
We are not claiming two sigma. Nobody should, on this evidence.
We do not predict anything. Every skill level in the record is derived from work a teacher marked and approved, and nothing in the product infers where a child will end up. If you changed a score, the record uses your score.
And a diagnosis is not a lesson. Knowing that six children are weak on paragraphing does not teach paragraphing to them. It tells you what Tuesday is for, which is a real improvement on not knowing, and considerably less than the "AI tutor for every child" framing implies.
The gap between one-to-many and one-to-one was never really about attention. It was about how long it takes to find out what each child cannot do yet, and how much preparation stands between knowing that and acting on it. That gap is closeable now, which is a smaller and more specific claim than two sigma, and one that holds up.
If you want the mechanics rather than the argument, personalised learning with AI: what it actually looks like walks through the loop from the classroom side, and the marking underneath it is covered in AI grading is only as good as the rubric behind it.