AI grading is only as good as the rubric behind it

AI grading is only as good as the rubric behind it

· 8 min read
Zippy Team
Zippy Team
Create, Grade and Personalize learning

The pitch for AI grading is usually volume: every child gets detailed feedback on every piece, instead of a tick and a number. It sounds unarguable. It is worth pausing on anyway, because the assumption underneath it is that more feedback is better feedback, and the evidence does not support that.

Kluger and DeNisi's 1996 meta-analysis is still the largest of its kind: 607 effect sizes across 23,663 observations. Feedback helped on average, at about d = 0.41. But in 38% of the experiments, feedback made performance worse. Not neutral. Worse.

So a tool that produces ten times more feedback is not obviously ten times more useful. It could be ten times more of the kind that backfires. The question worth asking about any AI marker is not how much it writes, but whether what it writes has the properties the research says feedback needs.

What the research says makes feedback work

Two findings do most of the practical work here.

The first is Kluger and DeNisi's own explanation for their result. Feedback interventions failed, they argued, when they moved a learner's attention away from the task and towards the self. "You are careless" is about the child. "This sentence has three ideas in it and needs to be two" is about the work. The first invites a defence, the second invites an edit. Teachers know this instinctively and it is the first thing that gets lost when marking is done at speed at 11pm.

The second is Hattie and Timperley's framework from The Power of Feedback (2007), which reduces effective feedback to three questions a learner should be able to answer afterwards:

  • Where am I going? What does good look like here?
  • How am I going? Where is my work against that?
  • Where to next? What is the specific thing I do about it?

Most marking answers the second question, sometimes. A number answers it very badly. The first and third are usually missing entirely, and they are the ones that make the second useful.

That framework is a decent specification for a grading tool, so it is the one we build against. Here is what each question turns into.

Where am I going: the rubric is the answer

Everything Zippy marks is marked against a rubric you control, not a general sense of good writing and not a scale we picked for you.

A score rubric is the familiar grid: your criteria down the side, bands across the top, and a descriptor in every cell saying what that band looks like. For a narrative you might use ideas and plot development, language and vocabulary, and organisation.

The thing that matters, and the thing most people underinvest in, is the descriptors rather than the labels. The descriptors are what the work is actually matched against. "Excellent" is not a specification; a sentence describing what an excellent piece of organisation looks like in a P5 narrative is. If your bands are landing differently from how you would band them yourself, that is a descriptor problem and it is fixable.

Where do the rubrics come from? Three ways, and they mix. Clone a ready-made set from Zippy Discover, which carries rubrics and skills lists aligned to real syllabuses. Adapt one, because a clone is yours and nothing is locked. Or build your own from an empty grid. Most departments start by cloning something close and rewriting the descriptors to match how they actually mark, which is much faster than authoring from nothing.

How am I going: tie it to the sentence, not to the child

This is where the task-versus-self distinction stops being a theory and becomes an interface decision.

Zippy returns a score per criterion, a comment, and inline annotations marked in place on the student's own words. The annotation is the part children actually read, and it is the part that keeps the feedback pointed at the writing. A note attached to a specific sentence is structurally incapable of being about the child's character. It is about that sentence.

That is also why we have resisted anything that scores effort, or produces an overall verdict on a student rather than on a piece. It reads as motivating and it is exactly the move Kluger and DeNisi found backfiring.

Where to next: this is what the skill map is for

A score tells you a child got 14 out of 20. It does not tell you which four skills cost them the marks, and that is the only version of the information you can plan a lesson around.

So Zippy grades against a second kind of rubric as well. A skills rubric lists the individual skills in a unit, each with levels describing what partial and full command look like. A narrative skills list might carry dozens, grouped by strand, with entries as specific as "show, not tell: sensory detail".

The two stack. A preset is what an activity actually points at, and it can carry a score rubric, a skills rubric, or both. When it carries both, one pass over a submission returns a mark against your criteria and a per-skill breakdown, plus the comment, the next step and the annotations.

That combination is the whole argument. A score you can put in a report, and underneath it a diagnosis you can teach from. Two children on the same 14 out of 20 rarely need the same lesson, and this is the view where you can see which one needs what.

Aligning to a standard without outsourcing your judgement

"Standards-aligned" is doing a lot of unexamined work in edtech copy at the moment, so it is worth being precise about what it can and cannot mean.

A rubric can be written to a published syllabus, and the descriptors can be lifted from the assessment objectives that syllabus sets out. That is real and it is useful, and it is what the Discover sets are. What it cannot mean is that a piece of software has privileged access to how a particular exam board would have marked a particular script. Nobody has that. Anyone claiming it is selling you a feeling of authority.

What you get from alignment is consistency and a shared vocabulary: every teacher in a department marking against the same descriptors, and a child hearing the same names for the same skills in September and in March. That is worth a great deal on its own, and it does not require anyone to pretend.

The twenty minutes that decide whether you can trust any of it

Here is the most useful thing in this post, and it takes one evening.

Take a set you have already marked by hand. Run it through the preset you plan to use. Compare the marks to your own, script by script.

You are not looking for identical numbers. You are looking for whether the disagreements have a pattern: if it is consistently a band harsher on organisation, that is a descriptor you can rewrite. Random disagreement and systematic disagreement mean completely different things, and twenty minutes on a set you already know the answers to is the only way to tell them apart.

Do this before you rely on it for anything that goes to a parent. We would rather you did it and found a problem than skipped it and found the problem later.

What we are not claiming

An AI marker inherits every weakness in your descriptors. Write a vague rubric and you will get vague marking applied with impressive consistency, which is arguably worse than vague marking applied inconsistently, because it looks authoritative.

Nothing goes to a student until you approve it. You can change a score, rewrite a comment, or throw the whole thing out for one child because you know something about that child that no rubric encodes. And nothing in the gradebook is inferred or predicted; it is all derived from marking you signed off.

We are not claiming a measured improvement in results. What the research supports is narrower and still worth having: feedback that names the success criteria, points at the work rather than the writer, and ends in a specific next step is the kind that tends to help rather than the kind that tends to backfire. Producing that reliably for thirty children, in an evening, is the actual product.

The mechanics of running a set are in marking a class set and how to mark a whole class without losing control. The step-by-step for rubrics themselves is in configuring rubrics.