Marking at Scale
Ninety scripts and a weekend. This is the situation automation was made for, and it is also the situation where the cost of automating badly is highest, because volume is exactly when you stop noticing what you have stopped noticing.
The goal is not to mark ninety scripts faster. It is to mark ninety scripts and still know at the end what your classes do not understand.
For comparison with workplace systems that formalise time or activity data, Monitask’s page on employee time clock software shows how the same type of measurement is handled outside education.
Sort before you start
The single highest-return move, and it takes ten minutes.
Read fifteen scripts — properly, unassisted, no marks awarded. Not the top fifteen or the ones on top of the pile; a spread.
For an external perspective on teaching, assessment, and feedback, see The Learning Scientists.
You are not marking. You are finding out what happened: the two or three errors that recur, whether the question was understood, which part of the task collapsed. That is the diagnosis, and it is the thing you actually need on Monday.
Everything after this is faster, because you now know what you are looking for, and because you have the picture before fatigue sets in. Marking the same fifteen scripts as numbers 76 to 90 would have told you nothing — by then you are pattern-matching, not reading.
Then split the pile
Mechanical scoring, where criteria are genuinely explicit, can be automated or done at speed. Spot-check for systematic bias by student, not just for accuracy.
Comment banks built from your fifteen-script diagnosis, not generic ones. You know the three errors; write three good responses to them once and reuse. This is the oldest efficiency in teaching and it is better than a generated equivalent, because it is anchored in what this cohort did.
One personal sentence each, written by you. Fifteen seconds per script, twenty-two minutes for ninety, and it is the part they will read.
A written summary for yourself, five lines, before you close the folder. What did they not get, what surprised you, what will you reteach. Write it the same day; it is gone by Tuesday.
What to protect at all costs
Never process the whole pile without reading any of it. If every script goes through a tool and you only see aggregates, you have marked nothing. You will get a report of common errors, and you will not get the thing you did not know to look for — which is where the interesting information always is.
Never let the aggregate be your only view. A summary tells you sixty percent made a particular error. It does not tell you that the six students who made a different error were all in the same group, or that one student's answer was wrong in a way that is genuinely more sophisticated than the correct one.
Never skip the fifteen. If you cut one thing under pressure, cut the depth of comments, not the reading.
Speed techniques that cost nothing
Independent of any tool, these survive scale.
Mark question by question, not script by script. You hold one mark scheme in your head, you are far more consistent, and comparison across students is immediate.
Mark in two passes. First pass fast, no comments, just a category. Second pass only on the boundaries, where your judgement actually changes an outcome.
Whole-class feedback instead of individual comments. For most tasks, a single page addressing the three common errors, given to everyone and worked through in class, does more than ninety individual notes. It is faster, and students act on it more because you talk them through it.
Stop writing what the mark already says. If the grade communicates the level, the comment should communicate the next step, and nothing else — one thing, doable, specific.
The failure mode with no alarm
Nothing goes wrong on the day. That is what makes this one dangerous.
The scripts get marked, the grades are defensible, the students get feedback. And over a term your lessons get slightly less well-aimed, because the picture of what your class understands is now assembled from summaries rather than from having read their work.
There is no moment where you notice. The symptom, if it appears at all, is a vague sense that a class is not where you thought — usually at the point where it is too late to fix cheaply.
The countermeasure is the fifteen scripts. It is the whole defence, it costs half an hour, and it is the first thing to go when the pile is large.
For departments
Two things worth agreeing rather than leaving to individuals.
A shared expectation about what gets read. If the department's position is that automated scoring is fine, say what sample every marker reads properly. Otherwise the conscientious carry the load and the exhausted do not, and nobody can say which.
A shared comment bank per assessment, built after the first cohort is marked. Written once by whoever marked first, used by everyone. This is a bigger saving than any tool and it improves consistency at the same time.
The short version
- Read fifteen scripts properly first, awarding no marks — that is the diagnosis and it is what you need on Monday
- Build comment banks from those fifteen rather than using generic ones
- One personal sentence each: fifteen seconds a script, twenty-two minutes for ninety
- Never let the aggregate be your only view; the interesting information is what you did not know to look for
- Question by question, two passes, whole-class feedback — all survive scale and cost nothing
- The failure has no alarm: lessons get less well-aimed over a term and nothing announces it