Rubrics That Survive Automation
Feed a rubric to a machine and you find out what your rubric actually says, as opposed to what you thought it said.
The discovery is usually uncomfortable. Words like insightful, sophisticated, perceptive and original are not criteria. They are placeholders — a shorthand between colleagues who already share a sense of what good work looks like in their subject, refined over years of moderation meetings and staffroom arguments.
For a workplace comparison of the underlying concept, Monitask’s discussion of interpersonal synchrony shows how the same behaviour is framed outside education.
A machine has none of that shared sense. Asked to apply "insightful," it does the only thing available: it looks for surface features that correlate with the word in the text it has seen. Length. Vocabulary. Structural conventionality. Hedged, balanced phrasing.
The three kinds of criterion
Sorting your rubric into these decides what can be automated.
For an external perspective on teaching, assessment, and feedback, see Carnegie Mellon Eberly Center.
Checkable. The thing is present or it is not. Three sources cited. A conclusion. Correct units. Paragraph structure. Machines do this reliably and it is tedious for humans.
Measurable. A quantity or a proportion. Word count, number of examples, percentage of steps completed correctly.
Judged. Requires knowing what good looks like in your subject and often in your year group. Insightful. Well-argued. Original. Perceptive. Appropriate for audience.
Most rubrics mix all three in one column without marking which is which, and then everyone is surprised when automated scoring produces confident nonsense on the third kind.
Rewriting a judged criterion
You cannot automate judgement, but you can often replace a placeholder with something observable that stands in for it. Not perfectly — usefully.
Insightful analysis becomes: identifies something not stated directly in the source and supports it with specific evidence from the text.
Well-argued becomes: states a position, gives at least two reasons, and addresses one objection.
Sophisticated vocabulary — delete it. This one is worth deleting outright rather than rewriting, and the reason matters: it is the same surface feature detectors use to accuse students, which means the same trait is rewarded in marking and punished in suspicion, and second-language students are hit at both ends.
Original becomes: makes a comparison or connection not made in class.
The rewrites are cruder than what you meant. That is the point of writing them down: a criterion you cannot express is a criterion you cannot moderate or teach, and students have been trying to guess at it for years.
What this does for students
Here is the genuine gain, and it is larger than the marking efficiency.
Students have always had to infer what "insightful" means from marks and comments. The ones who infer well are usually the ones with adults at home who know the code. Making the criterion explicit removes an advantage that had nothing to do with ability.
A student who is told "find something the source does not say directly and prove it from the text" can attempt that. A student told to be insightful can only try harder in an unspecified direction.
What to keep unautomated
Keep at least one criterion that is openly a matter of judgement, and label it as such.
Not everything worth rewarding can be specified in advance — that is why professional judgement exists rather than being an admission of vagueness. A rubric that is fully mechanical will be gamed within a term, because students optimise against whatever is written down, and a fully specified rubric is fully optimisable.
The workable arrangement is: mechanical criteria carry most of the mark and can be automated; one judged criterion is reserved, explained in your own words, and applied by you.
Practical steps
Take one rubric you already use. Mark each line C, M or J.
Run it against three pieces of work you have already marked, and compare the automated scores with your own.
Look at the disagreements, which are the whole exercise. Sometimes the machine caught something you missed. More often it has revealed that the criterion is a placeholder — and it will usually have scored the longest, most conventional piece highest.
Check the disagreement pattern by student. If the automated score is systematically lower for particular students, you have found a bias you can now name. Spot-checking for exactly this is not optional.
Rewrite the two worst lines. Not all of them. Two.
The by-product
Doing this makes you a better marker whether or not you automate anything.
Most rubrics are inherited, and most of us have never been forced to say precisely what we mean by the words in them. A machine taking them literally is an unusually blunt reviewer — and once the criterion is written properly, the feedback you give against it gets more specific too.
The short version
- "Insightful" and "sophisticated" are placeholders for shared professional judgement, not criteria
- Sort every line into checkable, measurable, or judged — only the first two automate
- Rewrite judged criteria into something observable, accepting that the rewrite is cruder than you meant
- Delete "sophisticated vocabulary": it is the same surface feature detectors use to accuse students
- Explicit criteria remove an advantage held by students who already know the code
- Keep one openly judged criterion, or the rubric will be fully optimised against within a term