Back to Blog AI in Education

How to Trust AI Marking: Mark the First Ten Yourself

If you are sceptical about letting AI near your marking, that instinct is doing its job. This is how to test it properly on a single class set, and why the teachers who try hardest to catch it out are usually the ones who end up trusting it.

ZH
Zubair Hasan
16 August 2026 · 8 min read

The worry is a reasonable one

I am not sure I would trust it. And honestly, I am not sure I want to. It feels like handing over the thing I am actually paid to do.

A head of maths said that to us last term, and it is the most useful sentence anyone has said to us about AI marking. We would rather start there than with a sales pitch, because it contains two distinct worries and they need different answers.

The first is about accuracy: will a machine get a child's grade wrong? The second is about the job: if the marking goes, what is left of it?

No vendor can settle the first one for you. It is settled by evidence you gather yourself, on your own scripts, in an afternoon, and this article will show you how. The second needs an honest account of what marking actually is, so take that one first.

What you would actually be handing over

Think about the last set of Year 11 mocks that sat on your desk.

Deciding whether a student earned the M1 when they have gone about it in a way the mark scheme never anticipated is professional judgement, and it took years to build. Reading someone's working and seeing that they understand the topic perfectly well but keep dropping a sign is teaching.

Writing "watch your sign when you expand the bracket" for the ninth time that evening is neither. Nor is transcribing thirty-one totals into a spreadsheet, or re-adding a column at half eleven because the numbers would not reconcile, or trying to turn a marked pile into anything resembling useful information about who needs what on Monday.

That second list is what a marking assistant takes off you. The first list stays exactly where it is.

You already have a word for this. It is second marking.

Every department already does this. Standardisation at the start of a series. Cross-marking a sample so that a grade means the same thing across three sets. Moderating coursework. Asking a colleague to glance at your borderline scripts before the data drop.

Nobody has ever walked out of a standardisation meeting feeling replaced. It is simply what professionals do with work that matters: they check it against someone else.

So treat Marky as a second marker. A fast one, with no view on who deserves a B, no memory of what a child scored last time, no soft spot for neat handwriting, and no fatigue at eleven o'clock on a Sunday night. It marks the same set you are marking, against the same mark scheme, and it shows its working: what it read on the page, which mark scheme point it matched that to, and how confident it was, question by question.

You remain the first marker. That is the entire model, and it is the model we would defend even if the technology got twice as good.

How to test it: try to catch it out

Do not take our word for the accuracy, and please do not take it on faith. Try to break it. Here is the protocol we would genuinely recommend, and it costs you one class set that you were going to mark anyway.

1. Pick a set you already have to mark

Not your hardest paper. Not your most consequential one. A routine mid-topic test is ideal. The whole point is that you lose nothing, because the work was on your desk either way.

2. Mark ten of them yourself, first

On paper, in your normal way, before you open anything on a screen. Write your marks down somewhere you cannot quietly revise them later. This is your control group, and it only works if it comes first.

3. Upload the set with its mark scheme, then open the same ten

Scans, a single class PDF that gets split into one script per student, or photographs taken on your phone. Whatever you already have. Marky reads the work, matches it against the scheme, and returns a mark per question with its reasoning attached. Your ten are sitting there waiting to be argued with.

4. Compare question by question, never just the totals

This is where informal trials usually go wrong. Two totals can agree by pure luck: a mark wrongly given on 3(a) and a mark wrongly withheld on 7(c) net out to a number that looks like perfect agreement. All of the real information lives at question level, which is where it should be reported to you.

5. Investigate every disagreement

This is the interesting part. The disagreements sort into three piles, and only one of them is about the software:

  • It misread the page. Cramped handwriting, a sprawling diagram, working that spills into the next answer box. Override it, which takes seconds, and make a note of it. You have just learned where its weak spot is.
  • The mark scheme is genuinely ambiguous, and the two of you read it differently. That is worth five minutes of your next department meeting whether or not you ever use AI marking again.
  • You would mark it differently now. This is the pile nobody expects, and experienced markers find it the most interesting. Marking drifts across a pile of thirty; script 4 and script 28 are not always seen by quite the same teacher. A second marker that treats script 28 exactly as it treated script 4 is doing something useful for your consistency, not just your workload.

Then do it again the following week, with the rest of the set or a different class. After two or three rounds you will have something worth far more than any vendor's accuracy claim: you will know exactly which questions you trust it on and which you do not.

That is calibrated trust rather than blind trust, and it is the only kind that survives contact with a real exam series. It is also why we would rather you started out sceptical.

Every correction you make, it keeps

Here is the part that changes the conversation, because it inverts the thing people are afraid of.

When you override a mark in Marky, that correction does not just fix one script. It is stored as an example of how you mark that kind of answer: the student's work, the mark the AI gave, the mark you gave, and why. The next time a question like it comes round, your correction is put in front of the marker as a worked exemplar before it decides anything.

The direction of travel matters. This is not a system training you to accept its judgement; it is a system being trained on yours. The more of your department's expertise goes into it, the closer it gets to marking the way your department marks.

Corrections travel sideways across the class as well. Adjust a mark on 6(b) by a mark or more and every other script in that set gets 6(b) flagged for you to spot-check, without changing anyone's score. If the marker got it wrong for one student, it may well have got it wrong for others, and you should be the one who decides.

Which is the honest reason to start now rather than in a year. The teachers who put their judgement in early are the ones who end up with a marker that thinks like them.

What we will not claim

Confidence built on overselling collapses the first time something goes wrong, and we would rather you found the limits from us than from a parent email. So, plainly:

It is not perfect, and we will not pretend otherwise

Measured against a verified examiner across 3,276 question sections, covering mathematics at GCSE and AS and physics at AS and A-level, Marky awards exactly the same mark on 92.4% of sections, comes within one mark on 98.9%, and within two marks on 99.8%. At AS and A-level, every section in that benchmark falls within two marks of the examiner. Your subject, your cohort and your scanner will each move those numbers, which is why the ten-script test above matters more than this paragraph does. The full method, the denominators and the limits of the evidence are set out on our accuracy page.

It is weakest where you would expect

Dense or faint handwriting, hand-drawn diagrams, and working that wanders across the page. It is strongest on structured, well-scanned scripts. None of that will surprise anyone who has marked a mock.

It tells you where to look rather than quietly guessing

Questions it is unsure about are badged for review with the reason attached. Whole scripts are pulled into a review queue when something warrants your eyes. That includes any paper landing within 5% of a grade boundary, where a single mark changes a grade. Your attention goes to the handful that need it instead of being spread evenly across the whole pile.

Your override always wins, and it is on the record

Change any mark and that is the mark: recorded, attributed and reflected everywhere downstream, with a full audit trail if a grade is ever disputed. There is no setting in which the software overrules a teacher.

Nothing reaches a student until you release it

Results are held back for your review by default, and you publish when you are ready, either a whole class or one script at a time. If a paper still has an open review flag or an ungraded question, Marky stops and tells you before it goes anywhere. No child, and no parent, sees a mark you have not looked at.

What comes back the other way

The hours are the obvious return, and we will not be coy about them. A set that eats a weekend becomes an evening's review instead.

But the return that surprises people is the information. When thirty scripts are marked consistently and recorded question by question, you can suddenly see things that are simply invisible in a pile of paper: that twenty-two of thirty-one dropped the same mark on 6(b); that your top set has a systematic problem with units rather than with the physics; that the student who slid a grade slid it entirely inside one topic and is fine everywhere else.

You see it the same evening the test was sat, while the reteach is still worth doing, rather than the following Monday when the lesson that would have fixed it has already gone.

That is not a smaller job. It is the same job with the tedious third removed and much better evidence bolted on.

And the job question, directly

It deserves a straight answer rather than a reassuring one.

Marking is not what makes you a teacher. It is a proxy, an expensive and exhausting one, for two things that matter: knowing where each student is, and telling them something useful about it. Students learn from the feedback, not from the act of you physically writing it out thirty-one times. The scarce resource in any school is teacher attention, and marking has always consumed an enormous quantity of it while returning very little per hour spent.

We will also say this plainly. A school that cuts teaching posts because marking got faster has not been disrupted by technology; it has made a decision about what it values, and it could have made that decision at any point in the last twenty years. The far likelier outcome is that the recovered hours go back into what marking was crowding out. Intervention. Actual feedback conversations. Planning something better than you had time for last year. Leaving the building at a sensible hour.

Nobody is arguing that a machine should decide what a child is capable of. The argument is narrower than that. It can transcribe, cross-reference a mark scheme, add up and organise, so that the person who can judge what a child is capable of has the time and the evidence to do it properly.

Try one set. That is the whole ask.

You do not have to decide anything today. You do not have to change how your department marks, or announce it to anyone, or commit to a term of it.

Pick one class set you were going to mark anyway. Mark ten of them yourself, properly, the way you always do. Then upload the set and see how it compares, question by question.

If it is not good enough, you will know inside an hour, and you will have marked ten scripts you had to mark regardless. Nothing lost, and you will have a far sharper view of what this technology can and cannot do than any article could give you.

And if it is good enough, you have got your Sunday back.

Start your free 3-day trial

Frequently asked questions

Will AI marking replace teachers?

No. AI marking automates transcription, mark-scheme cross-referencing and totalling. It does not automate professional judgement or teaching. Every mark remains editable by the teacher, and the teacher decides what a result means and what happens next. The realistic effect is on workload, not on headcount.

How do I check whether an AI marker is accurate enough for my classes?

Mark ten scripts yourself first, on paper, then run the same set through the tool and compare question by question rather than by total. Totals can agree by coincidence when two errors cancel out. Repeat over two or three sets and you will know which question types you can trust it on.

What happens when I disagree with a mark?

You override it, and your mark is the one that stands, recorded with an audit trail. Your correction is also kept as an exemplar of how you mark that kind of answer, and if the adjustment is large enough the same question is flagged for spot-check on the rest of the class without changing anyone's score.

Can students see results before I have checked them?

No. Results are held back for teacher review by default and are only visible to students and linked parents once you publish them, either as a whole class or one script at a time.