
How AI helps teachers mark PSLE compositions
Thirty compositions to mark, and each one needs a Content mark out of 18, a Language and Organisation mark out of 18, and a comment the child can actually do something with.
What usually happens is that the marking gets thinner towards the bottom of the pile. The comments get shorter, the next steps get vaguer, and the last few children get less than the first few did. By then it is eleven at night, which is most of the explanation.
Here is what changes when the first pass is already done before you sit down.
It marks against your bands, not a generic scale
This is the part most tools get wrong. A marker carrying its own idea of good writing will give you marks you have to argue with on every script.
Zippy holds the PSLE P5/6 Continuous Writing rubric as a rubric: Content and Language & Organisation, six levels each from No Credit to Exemplary, with the mark range attached to every band.

The bands as Zippy holds them. If your centre marks to its own descriptors, you replace these with yours and it marks to those instead.
What each of those bands actually asks for, and two real scripts marked against them, is covered in how a PSLE composition is marked. This post is about what happens when you point it at a class set.
What comes back for each script
Two marks, not one total. A composition sitting at Content 15 and Language 9 needs a completely different conversation from one at Content 9 and Language 15, and a single number out of 36 hides that.
Underneath each mark, the reasoning: which band it landed in, which descriptor it matched, and which line in the child's writing put it there. Annotations sit on the script itself, each one tagged with the criterion it counts towards, so a comment about tenses is visibly a Language comment rather than a general complaint.
That trail is what makes a mark arguable. You can disagree with it specifically, on the line, rather than in general.
It reads handwriting from a photograph
Most PSLE composition practice is still done on paper. Tools that only accept typed text quietly hand you a second job first, and typing thirty compositions is worse than marking them.
Photograph the pile. Zippy transcribes each script and marks the transcription. A phone camera on a page of P5 handwriting is enough.
The honest limit: genuinely difficult handwriting produces transcription errors, and a transcription error becomes a marking error. The scripts most likely to be misread are the weakest ones, which is the wrong way round. Those are the ones to read yourself.
How often it agrees with you
Researchers at the University of Georgia measured how often an AI model landed on the same mark as a human marker. It was 33.5%. Give the model the teacher's own rubric and separate work puts it above 50%.
Neither number is a marking tool you can leave alone, and we would rather say so. What the first pass buys is the hour of reading thirty pieces and writing thirty sets of comments from nothing. The mark itself is still a judgement, and it stays yours.
That gap between a third and a half is also the argument for the section above. A marker holding your descriptors does not have to guess at what your centre counts as "sufficiently developed".
Nothing goes out until you release it
Not the marks, not the annotations, not the feedback. A graded class set sits unreleased until you say otherwise, and while it is held, the child's own view has nothing in it.
Every comment is yours to reword or delete, every mark is yours to change, and what you change survives the next run rather than being overwritten by it. The mechanics of that, and why the "it forgot what I decided" problem is the one that kills AI markers, are in how to mark with AI and still be the one who marked it.
What reaches a parent is the mark and the comment under your name, which is what a report has always been.
Thirty at a time, and the standard holding
Speed is the obvious benefit. Consistency is the one that matters more, and it is harder to see.
The twenty-eighth script is read by a different person than the first. Not a careless person. A tired one, at eleven at night, whose sense of what "sufficiently developed" means has drifted a little since script three. Across four teachers in a centre, that drift becomes a fairness problem, and parents notice it before anyone internally does.
A rubric applied by machine does not drift. It puts the same descriptor against script 30 as against script 1 and shows its reasoning both times, so a head of department can pull two scripts marked a week apart and see whether the standard held.
You still decide. You are deciding against a consistent first pass rather than against your own memory of what you did on Tuesday.
Where it struggles
It does not know your class. It cannot know that this child has written the same ending all term, or that this one reached 400 words for the first time. The band can be right while the context is missing, and the context is often the more useful half.
Voice is invisible to a rubric, so it is invisible here. A composition can be funny, strange and completely alive, and the band descriptors have no vocabulary for it. That judgement belongs to the person who knows what the child is capable of.
It will not save the evening if you read all thirty in full. Teachers who spot-check the bands they distrust save time. Teachers who re-read everything save very little. That is a real trade-off, and better said now than discovered in week two.
Try it on one you have already marked
The honest test is a script whose mark you already know. Run it, and see where it disagrees with you and where it does not.
The composition grader on the grading page takes a single script without an account. Zippy is free for teachers, and the free tier is a real tier rather than a trial.
AI and human marker agreement, and the effect of supplying a rubric: arXiv 2504.13557.