Marking and Feedback Are Not the Same
A stack of work lands on your desk and you call the next two hours "marking." Inside those two hours are three different jobs, and they have almost nothing in common except arriving together.
Scoring. Deciding what this piece is worth against a standard.
For comparison with workplace systems that formalise time or activity data, Monitask’s page on time tracking software shows how the same type of measurement is handled outside education.
Responding. Telling this student something that changes what they do next.
Diagnosing. Working out what the class has and has not understood, which determines the next lesson.
For an external perspective on teaching, assessment, and feedback, see Visible Learning.
They get automated as a unit because they feel like one task. That is the mistake, because their answers are completely different: one machines well, one machines partially, and one should not be delegated at all.
Scoring: machines well, within limits
Applying a fixed standard to a piece of work is the most mechanical of the three, and it is where automation is least controversial.
It works where the criteria are genuinely explicit and the judgement is close to matching: right or wrong answers, checkable structure, presence or absence of required elements, mechanical accuracy.
It works badly where the rubric contains words like insightful, sophisticated, or original. Those are not criteria, they are placeholders for a professional judgement, and a tool asked to apply them will produce a confident number derived from surface features — usually length, vocabulary, and structural conventionality.
Which is worth knowing for a specific reason. Those are also the features that make a detector flag a student. A system that rewards elaborate prose in scoring and suspects plain prose in detection is applying the same crude proxy twice, in opposite directions, to the same child.
If you automate scoring, spot-check. Not one in a hundred — enough that you would notice a systematic bias, which usually means enough to see whether the same kinds of students are being scored down.
Responding: machines partially
A first draft of feedback is a genuine use, and it is where most of the measured time saving comes from.
What machines produce well: a clear statement of what the criteria say, comments on structure and mechanics, a suggestion for a next step in generic terms.
What they cannot produce, because they lack the information: anything that depends on knowing this student. That you told them this exact thing three weeks ago. That this is the first time they have attempted an argument at all and the attempt is the achievement. That they are capable of much more and are not trying. That something is going on at home.
A student can tell the difference instantly, and the cost of getting caught is high. Feedback that reads as generic teaches a student that nobody read their work, which is worse than no feedback, because it converts effort into evidence of being ignored.
The workable arrangement is a draft you then make specific. If you find yourself accepting drafts unedited because they read well enough, the tool has quietly replaced the response rather than started it — and that drift has a shape and a tell.
Diagnosing: do not delegate
This is the one that matters and the one nobody defends, because it is invisible in the output.
While you work through a stack, you build a picture: six students made the same error, which means the explanation on Tuesday did not land. Two of them made a much more interesting error, which means they understood something and applied it wrongly. Someone who normally coasts has suddenly engaged. Someone who normally engages has stopped.
None of this is in the marks. It is a by-product of the attention, and it is most of what makes the next lesson good rather than merely planned.
A summary generated from the work will give you the aggregate — common errors, distribution of scores. It will not give you the thing you did not know you were looking for, because it can only report against categories it was asked about.
So: if you automate the first two jobs, you have to buy the third back deliberately. Read a sample yourself, unassisted, with the specific intention of noticing rather than scoring. Write down what you noticed, because the picture fades within a day and it is the part you will actually use.
The practical split
For a typical stack:
Score mechanically where the criteria are genuinely explicit, and spot-check for systematic bias.
Draft responses, then make each one specific to the student — one sentence of genuinely personal comment is worth a paragraph of accurate generic advice.
Read a real sample yourself, unassisted, to diagnose. Ten pieces of work read properly tell you more about the class than fifty processed quickly — and at ninety scripts this is the first thing to go and the last thing you should cut.
Notice what you are no longer noticing. This is the failure mode with no alarm attached. Nothing goes wrong on the day; the lessons just get slightly less well-aimed over a term, and there is no moment where you can see it happening.
The question to keep asking
For each part of the pile: am I doing this to produce a mark, to help this student, or to find out what happened?
The first can mostly be handed over. The second can be started and must be finished by you. The third is the reason marking is part of teaching rather than clerical work, and giving it away costs you something you will not notice losing.
The short version
- Three jobs arrive as one pile: scoring, responding, diagnosing
- Scoring automates where criteria are explicit; words like "insightful" are placeholders for judgement, not criteria
- Automated scoring rewards elaborate prose — the same crude proxy detectors use to accuse students
- Draft responses are fine; unedited drafts teach a student that nobody read their work
- Diagnosis is a by-product of attention and cannot be delegated; buy it back deliberately with a read sample
- The failure mode has no alarm: lessons just get slightly less well-aimed over a term