The worry is a reasonable one
I am not sure I would trust it. And honestly, I am not sure I want to. It feels like handing over the thing I am actually paid to do.
A head of maths said that to us last term, and it is the most useful sentence anyone has said to us about AI marking. It contains two distinct worries, and they need different answers.
The first is about accuracy: will a machine get a child’s grade wrong? The second is about the job: if the marking goes, what is left of it?
No vendor can settle the first for you. It is settled by evidence you gather yourself, on your own scripts, in an afternoon, and this article shows you how. The second needs an honest account of what marking actually is, so take that one first.
What you would actually be handing over
Deciding whether a student earned the M1 when they went about it in a way the mark scheme never anticipated is professional judgement. Reading someone’s working and seeing that they understand the topic perfectly well but keep dropping a sign is teaching.
Writing “watch your sign when you expand the bracket” for the ninth time that evening is neither. Nor is transcribing thirty-one totals into a spreadsheet, or re-adding a column at half eleven, or trying to turn a marked pile into useful information about who needs what on Monday.
The second list is what a marking assistant takes. The first stays exactly where it is.
The roles swap. You become the second marker.
The comfortable version of this article says you remain the first marker and the software checks your work afterwards. A safety net. It is the easier thing to sell and it does not survive five minutes of arithmetic: if you have already marked all thirty-one scripts, the weekend has gone. A second opinion on finished work earns its keep at standardisation or on a disputed grade, but not set after set, and we would rather say so than sell it to you.
So, plainly: Marky marks first, and you take the second marker’s chair.
Every department already runs a two-marker model. Standardisation at the start of a series. Cross-marking a sample so a grade means the same thing across three sets. Moderating coursework. Nobody has ever walked out of a standardisation meeting feeling replaced, and notice who does what in that model: the person grinding through the first pass is not the senior one in the room. Second marking is the job you get promoted into.
Marky suits a first pass and nothing else. No view on who deserves a B, no memory of what a child scored last time, no soft spot for neat handwriting, no fatigue at eleven on a Sunday night. It marks against the scheme and shows its working question by question: what it read on the page, how it applied the mark scheme to it, how confident it was.
Then you moderate. You look at what it flagged, sample beyond that, overrule it wherever it is wrong, and decide what the results mean for Monday. Your mark is the mark, and nothing reaches a student until you release it.
How to test it: try to catch it out
Do not take the accuracy on faith. Try to break it. Here is the protocol, run once, on one class set you were going to mark anyway.
1. Pick a set you already have to mark
Not your hardest paper, not your most consequential one. A routine mid-topic test is ideal, because you lose nothing: the work was on your desk either way.
2. Mark ten of them yourself, first
On paper, in your normal way, before you open anything on a screen. Write your marks somewhere you cannot quietly revise them later. This is your control group, and it only works if it comes first.
3. Upload those ten with their mark scheme
Scans, a class PDF split into one script per student, or phone photographs. Marky returns a mark per question with its reasoning attached, and your ten are sitting there waiting to be argued with.
4. Compare question by question, never just the totals
This is where informal trials go wrong. Two totals can agree by pure luck: a mark wrongly given on 3(a) and a mark wrongly withheld on 7(c) net out to what looks like perfect agreement. All the real information lives at question level.
5. Investigate every disagreement
The disagreements sort into three piles, and only one is about the software:
- It misread the page. Cramped handwriting, a sprawling diagram, working that spills into the next answer box. Override it, which takes seconds, and note it. You have just learned where its weak spot is.
- The mark scheme is genuinely ambiguous, and the two of you read it differently. Worth five minutes of your next department meeting whether or not you ever use AI marking again.
- You would mark it differently now. The pile nobody expects. Marking drifts across a pile of thirty, and script 4 and script 28 are not always seen by quite the same teacher. A marker that treats them identically is doing something for your consistency, not just your workload.
One set settles it. Ten scripts compared question by question is a few hundred marking decisions, far more evidence than any vendor’s accuracy claim. You will know which questions you trust it on and which you do not. That is calibrated trust rather than blind trust, and it is the only kind that survives an exam series.
Be clear about what this is, though. It is an audit, not the arrangement. You mark ten scripts twice so that you never have to mark thirty twice again.
What the week after looks like
Say the audit went well. Here is Sunday from then on, which is the part most articles about AI marking leave out entirely.
- Upload the set. Scans from the department copier, a class PDF split into one script per student, or photographs off your phone. It comes back marked question by question, reasoning attached to every mark.
- Open the review queue, not the pile. Questions where the marker’s own confidence was low are badged as such, and the class set comes back ranked by how close each total sits to a grade boundary, closest first. Work down that list in the time you have.
- Sample, the way you already moderate. Read three or four scripts you were not directed to, one from the top of the set and one from the bottom. Not to re-mark them: to catch the one failure mode nothing can flag for you, which is a marker being confidently wrong without knowing it.
- Read the class breakdown. This is the part that changes Monday rather than Sunday.
That is an evening, and the last step is something you did not previously have at all.
Re-run the audit when the ground moves. Trust is per question type, not permanent. A new paper, a new subject, a switch to phone photographs, a cohort with markedly worse handwriting: mark ten yourself again. Any tool that discourages you from re-checking it is the wrong tool.
A correction can travel across the class
When you override a mark, you can ask Marky to show you every other script in the set where it made the same decision on that question, with exactly what would change on each. Nothing moves until you say so, and anything you apply can be put back. If it got 6(b) wrong for one student it may have got it wrong for others, and you should be the one who decides.
The direction of travel matters. This is not a system training you to accept its judgement; it is a system taking correction from yours.
What we will not claim
Letting it mark first is a bigger ask than a second opinion, and it obliges us to be straight about where it falls down. We would rather you found the limits from us than from a parent email.
It is not perfect, and we will not pretend otherwise
We compared Marky against the marks a class teacher had already written on their own scripts: 1,830 question sections across 60 real GCSE maths scripts, with nobody adjudicating between the teacher and us. Marky awards exactly the same mark on 90.8% of sections, comes within one mark on 99.2% and within two on 99.8%. Where it differs it is usually the more generous, by a single mark, which over a whole script comes to about 1.5 marks above the teacher’s total. We also measured an AS cohort, reported separately because that teacher could see Marky’s mark while marking. Your subject, your cohort and your scanner will each move those numbers, which is why the audit above matters more than this paragraph. Full method and denominators are on our accuracy page.
It is weakest where you would expect
Dense or faint handwriting, hand-drawn diagrams, working that wanders across the page. It is strongest on structured, well-scanned scripts. None of that will surprise anyone who has marked a mock.
It tells you where to look rather than quietly guessing
Unsure questions are badged for review, and the class set comes back ranked by how close each total sits to a grade boundary, closest first. Grade boundaries are roughly ten percentage points apart, so most of a set is within five of one: what carries the information is the order, not the flag. You work down the list in the time you have, and the papers where a single mark moves the reported grade are the ones at the top. This is the mechanism the whole arrangement rests on, which is why we would rather you tested it than believed it.
Your override always wins, and it is on the record
Change any mark and that is the mark: recorded, attributed, reflected everywhere downstream, with an audit trail if a grade is ever disputed. There is no setting in which the software overrules a teacher.
Nothing reaches a student until you release it
Results are held for your review by default and you publish when ready, a whole class or one script at a time. If a paper still has an open review flag or an ungraded question, Marky stops and tells you first. No child, and no parent, sees a mark you have not looked at.
And the job question, directly
Yes, this replaces something, and it is worth being precise about which word. It replaces the first pass: the transcribing, the cross-referencing, the totalling, the ninth writing-out of the same correction. It does not replace the marking decision, which stays yours, editable at any point and recorded as yours. Those two things get called by the same name, which is how this argument usually goes wrong.
Marking is not what makes you a teacher. It is a proxy, an expensive and exhausting one, for two things that matter: knowing where each student is, and telling them something useful about it. Students learn from the feedback, not from the act of you writing it out thirty-one times.
What you get back is not only the hours. When thirty scripts are marked consistently and recorded question by question, you see things that are invisible in a pile of paper:
- twenty-two of thirty dropped the same mark on 6(b);
- your top set has a problem with units rather than with the physics;
- the student who slid a grade slid it entirely inside one topic.
And you see it the same evening the test was sat, while the reteach is still worth doing.
And plainly, on the fear underneath the question: a school that cuts teaching posts because marking got faster has not been disrupted by technology. It has made a decision about what it values, and it could have made that decision at any point in the last twenty years. The likelier outcome is that the recovered hours go back into what marking was crowding out. Intervention. Actual feedback conversations. Leaving the building at a sensible hour.
If you are running that afternoon against more than one product, the questions worth asking each of them, including the ones we answer badly, are set out separately.
Try one set. That is the whole ask.
You are not deciding anything today, and you would not be deciding for your department in any case. Marky is bought per department rather than per teacher, so nothing changes for your colleagues until your colleagues agree it should. What one teacher can do in an afternoon is settle the question that meeting would otherwise argue about with nothing but a vendor’s figures on the table.
Pick one class set you were going to mark anyway. Mark ten of them yourself, properly, the way you always do. Then upload those ten and see how they compare, question by question. The whole exercise fits inside the 14-day trial, which is deliberate: a test you have to buy your way through to finish is not a test.
If it is not good enough, you will know inside an hour, and you will have marked ten scripts you had to mark regardless. Nothing lost.
And if it is good enough, you walk into that meeting with question-level agreement figures from your own scripts, in your own subject, marked to your own standard. That is a considerably better argument than anything on this website.