AI marking accuracy, measured against a verified examiner
A claim about marking accuracy is only as good as the benchmark behind it. This page sets out the figures, the denominators they are drawn from, the method that produced them, and the boundary of what the evidence supports.
The measured figures
Every figure below is drawn from the same benchmark of 3,276 question sections marked by a verified examiner, spanning mathematics at GCSE and AS, and physics at AS and A-level.
no difference at all, section by section.
a difference of at most a single mark in either direction.
of the mark the examiner awarded for the same section.
the denominator behind every percentage on this page. Each figure includes the one above it, so a wider tolerance can only ever count more sections, never fewer.
At AS and A-level, every section in the benchmark falls within two marks of the examiner. Percentages are published with their denominators because a percentage without one is not evidence.
How the accuracy figures are counted
Marking accuracy is measured at the level of the question section, not the paper. A section is the smallest unit the mark scheme awards against, such as question 4(b)(ii), and it is the unit a teacher would argue about in a moderation meeting. Comparing whole-paper totals would hide compensating errors, where a mark lost on one question is quietly returned on another, so the comparison is made section by section and the results are then aggregated.
Marks that match exactly
A section is exact when the mark awarded is identical to the examiner's mark for that section. This is the strictest of the three measures and the least forgiving, because a single mark of difference on a four-mark question counts as a failure even though the reported grade would often be unchanged. Across the benchmark, 92.4% of sections are exact.
Within one mark and within two marks
The tolerance measures record how far the difference extends when a section is not exact. Within one mark covers every section where the difference is at most a single mark in either direction, and stands at 98.9%. Within two marks covers every section where the difference is at most two, and stands at 99.8%. Both tolerances are published rather than a single headline number, because the distance of a disagreement matters as much as how often one occurs.
Why both are reported
Reporting only the exact figure understates the practical agreement between the two markers, since most of the remaining sections differ by a single mark on a long method question where two competent examiners would also differ. Reporting only the tolerance figure overstates it, because a generous tolerance can conceal a systematic tendency to award or withhold a mark. Publishing both, with the denominator attached, is the only presentation that allows a head of department to judge the size and the shape of the disagreement rather than accept a number.
The benchmark: real scripts, marked by a verified examiner
The comparison set is made up of real student scripts, handwritten under exam conditions, with the difficulties that implies: crossed-out working, answers continued in the margin, method that arrives at the right answer by an unexpected route. Each paper was marked by a verified examiner against the official mark scheme, and that marking is treated as the reference. The same papers were then marked by Marky, and the two sets of marks were compared section by section.
The benchmark covers four cohorts. Section counts differ between them, and the smallest cohort carries the least weight, so the per-cohort results are published in full rather than folded into a single average.
| Cohort | Question sections | Exactly the same mark |
|---|---|---|
| GCSE mathematics | 1,904 | 91.75% |
| AS mathematics | 638 | 94.67% |
| A-level physics | 668 | 92.37% |
| AS physics | 66 | 87.88% |
| All cohorts | 3,276 | 92.4% |
AS physics is the smallest cohort at 66 sections. A figure drawn from 66 sections moves by more than one and a half percentage points if a single section changes, so it is reported for completeness and should not be read as a stable estimate.
Two independent marking passes
Every paper is marked twice. The second pass re-reads the paper against the mark scheme and challenges the marks the first pass awarded, rather than confirming them. Where the two passes disagree, the disagreement is resolved before the mark is released. No mark reaches a teacher with that disagreement still open.
The reasoning behind each award stays attached to the question. A mark that arrives without a justification cannot be checked, and a marking system that cannot be checked cannot be trusted with a reported grade. Because the reasoning is recorded at section level, a teacher reviewing a paper can read why a mark point was given or withheld and form a view in seconds rather than re-marking the question from the beginning. The how it works page sets out the full sequence from upload to returned results.
Targeted review, with the reason stated
Papers that warrant human attention are flagged, and each flag carries one of six stated reasons. The reason is shown alongside the paper, so the review starts from a stated concern rather than from scratch.
Near a grade boundary
A total sitting close to a boundary is where a single mark changes the reported grade, so the mark carries more consequence than its size suggests.
Handwriting not read confidently
Where the handwriting could not be read with confidence, the paper is flagged rather than marked on a guess at what the student wrote.
Possible integrity concern
Raised for a teacher to consider. A flag of this kind is advisory and does not alter the mark on its own.
Disagreement between passes
The two marking passes reached different conclusions on the same section, which is a direct signal that the question is difficult to mark.
Arbitration
The disagreement required a resolution step before the mark could be released, and the paper is shown so that the resolution can be seen.
Wide variance
The spread of evidence for the section was unusually wide, which is treated as a reason to look rather than a reason to adjust.
Why flagging a subset beats re-checking everything
A department asked to re-check every paper has gained nothing: the marking has been done twice, and the second reading is the one that will be rushed. Attention is finite, and spreading it evenly across thirty papers guarantees that the two that needed it received the same thirty seconds as the twenty-eight that did not. Flagging a subset concentrates that attention where the evidence says the mark is least secure, and states the reason so the teacher knows what to look at. The remainder can be signed off as returned. It is the logic of moderating a sample chosen for a reason, rather than marking every paper a second time.
The teacher holds the final mark
Any mark can be changed by the teacher before results are published, on any question, with or without a flag. Nothing is locked, and no result reaches a student until the teacher releases it. The marking is a first draft produced quickly and consistently, and professional judgement remains where it belongs.
This matters for the way accuracy should be read. The figures above describe the quality of the draft, not the quality of the result a class receives, because the result a class receives has passed through a teacher who can and does correct it. A department that overrides a handful of marks per class set is using the system exactly as intended.
The limits of this evidence
A number is only believable if the boundary around it is stated. Three limits apply to everything above.
The measurement covers mathematics and physics
The 3,276 sections are mathematics at GCSE and AS, and physics at AS and A-level. Marky also marks biology and chemistry, and those subjects are not represented in this benchmark. No claim is made here about accuracy on them, and any figure quoted for them would be an extrapolation rather than a measurement.
Handwriting quality affects reading
Marking begins with reading, and a paper that a human examiner would struggle over is a paper the system will struggle over as well. Faint pencil, heavy crossing out, working that runs across two pages and a photograph taken at an angle all make a section harder to read correctly. Where reading confidence is low the paper is flagged, but the effect on accuracy is real and varies with the class and the scan.
The benchmark is a specific set of papers
These are particular papers, from particular specifications, marked by a particular verified examiner. They are not a random sample of every paper sat in the country, and the figures should be read as a description of that set rather than a universal guarantee. A different set of papers would produce a different number.
Stating these limits costs nothing that was ever true. It is also the reason the figures can be relied on: a benchmark that admits what it does not cover is a benchmark that has not been selected to flatter. The same reasoning appears in the comparison with manual marking, where the honest baseline is not perfect marking but tired marking on a Sunday evening.
What is marked, and how quickly
Marking covers mathematics, biology, chemistry and physics at GCSE, AS and A-level, against official mark schemes from AQA, Pearson Edexcel, OCR, WJEC/Eduqas and CCEA. Papers are accepted as PDFs and as photographs in .jpg, .png, .heic or .webp with no conversion required, so a set can be captured on a phone at the end of a lesson.
Full details of allowances and top-ups are set out on the pricing page, and the marking features that sit on top of the results are listed under features.
Questions departments ask about accuracy
What counts as an exact match?
An exact match means the mark awarded for a single question section is identical to the mark a verified examiner awarded for that section. It is measured per section rather than per paper so that compensating errors cannot cancel out.
How large is the benchmark?
3,276 question sections across four cohorts: GCSE mathematics with 1,904 sections, AS mathematics with 638, A-level physics with 668 and AS physics with 66.
Does every paper have to be checked by a teacher?
No. Papers warranting attention are flagged with one of six stated reasons, and the remainder can be signed off as returned. Any paper can still be opened and reviewed at will.
Can a teacher change a mark?
Yes. Any mark on any question can be overridden before results are published, and nothing reaches a student until the teacher releases it.
Which subjects were measured?
Mathematics at GCSE and AS, and physics at AS and A-level. Biology and chemistry are marked but are not covered by this benchmark.
The benchmark that counts is the department's own
The strongest test of any of this is not a figure on a marketing page. During the free 3-day trial, submit a set of papers that has already been marked in the department and compare the results question by question, against the specification taught and the handwriting produced. Agreement on a familiar class set settles the question in a way no published percentage can. No card is required.