AI marking accuracy, measured against the marks teachers actually gave
Marky awarded exactly the mark the class teacher gave on 90.8% of 1,830 GCSE maths question sections, and came within one mark on 99.2%. Those are 60 real scripts, marked in red pen by the teacher who taught the class before Marky saw them, with nobody adjudicating between the two. An AS cohort is reported separately further down, because that teacher could see Marky’s mark while marking. Measured against our reading of the published mark scheme rather than against each other, Marky is the closer of the two markers: 96.1% to the teacher’s 93.1% on the GCSE cohorts, and 96.5% to 94.0% on AS. On physics, where we read every section against the mark scheme ourselves, Marky matches our reading on 93.2% of 605 sections. This page gives the figures, the denominators behind them, the method that produced them and the boundary of what the evidence supports.
The measured figures
Every figure below is drawn from the same comparison: 1,830 question sections on 60 real GCSE maths scripts, marked by the class teacher before Marky saw them, and then by Marky. The AS cohort is not in these four figures; it has its own row in the table further down.
no difference at all, section by section.
a difference of at most a single mark in either direction.
of the mark the teacher wrote for the same section. Three sections in 1,830 differ by more.
the denominator behind these four figures, from 60 GCSE scripts. Each figure includes the one above it, so a wider tolerance can only ever count more sections, never fewer.
Nobody adjudicated these figures. There is no separate “correct” mark that Marky and the teacher were both scored against. Agreement means agreement with the mark the teacher wrote, and every difference is counted against Marky by default. Percentages are published with their denominators because a percentage without one is not evidence.
How the accuracy figures are counted
Marking accuracy is measured at the level of the question section, not the paper. A section is the smallest unit the mark scheme awards against, such as question 4(b)(ii), and it is the unit a teacher would argue about in a moderation meeting. Comparing whole-paper totals would hide compensating errors, where a mark lost on one question is quietly returned on another, so the comparison is made section by section and the results are then aggregated.
Marks that match exactly
A section is exact when the mark Marky awards is identical to the mark the teacher wrote for that section. This is the strictest of the three measures and the least forgiving, because a single mark of difference on a four-mark question counts as a failure even though the reported grade would often be unchanged. Across the 1,830 GCSE sections, 90.8% are exact.
Within one mark and within two marks
The tolerance measures record how far the difference extends when a section is not exact. Within one mark covers every section where the difference is at most a single mark in either direction, and on the GCSE sections stands at 99.2%. Within two marks covers every section where the difference is at most two, and stands at 99.8%. Both tolerances are published rather than a single headline number, because the distance of a disagreement matters as much as how often one occurs. The full distribution, including the direction of each difference, is set out below rather than summarised.
Why both are reported
Report only the exact figure and you understate the agreement, because most of the sections that differ are a single mark on a long method question. Report only the within-one figure and you overstate it, because a generous tolerance can hide a steady tendency to award or withhold a mark. In this comparison there is one, which is why the direction of every disagreement is published below. Both figures, with the denominator attached, let a head of department judge the size and the shape of the disagreement rather than accept a number.
The comparison set: real scripts, marked by the class teacher
The comparison set is made up of real student scripts, handwritten under exam conditions, with the difficulties that implies: crossed-out working, answers continued in the margin, method that arrives at the right answer by an unexpected route. The GCSE scripts had already been marked by the teacher who taught the class, in red pen, before Marky ever saw them; the AS teacher’s marks were supplied afterwards on a spreadsheet, which the limits section below explains. That marking is the reference. The same scripts were then marked by Marky and the two sets of marks were compared section by section, with the teacher's mark counted as correct in every case of disagreement.
Nobody was asked to decide who was right. There is no third mark, no adjudication step and no appeal: a section counts as exact only where Marky reproduced the teacher's own number. That is a harder test than it sounds, and a deliberately conservative one.
Three cohorts make up the set. Section counts differ between them, so the per-cohort results are published in full rather than folded into a single average. The headline is the two GCSE papers together. The AS row, and the row for all three cohorts, are shown in full but are not the figure to rely on, for the reason given under the table.
| Cohort | Scripts | Question sections | Exactly the teacher's mark | Within one | Within two |
|---|---|---|---|---|---|
| GCSE mathematics, Paper 1 (Higher) | 30 | 900 | 90.8% | 99.2% | 99.9% |
| GCSE mathematics, Paper 2 (Higher) | 30 | 930 | 90.8% | 99.2% | 99.8% |
| Both GCSE papers, the headline | 60 | 1,830 | 90.8% | 99.2% | 99.8% |
| AS mathematics | 18 | 720 | 90.7% | 99.2% | 99.9% |
| All three cohorts, including AS | 78 | 2,550 | 90.7% | 99.2% | 99.8% |
The three cohorts land within 0.1 of a percentage point of each other, on three different papers at two different levels. A single re-run of the same papers moves each cohort by about half a point, so the closeness is partly chance; what matters is that all three sit near 90%.
The GCSE rows are the headline because those teacher marks were read off the teacher's own ink on the physical scripts, with no software mark anywhere in view. The AS sheet was different. It showed Marky's proposed mark next to the empty column the teacher was filling in. Two things say they were marking rather than copying. They disagreed with the mark in front of them on 63 of the 720 sections. And their marks added up to a separate list of paper totals on all 18 scripts. Even so, a teacher who can see a mark may be nudged toward it, so AS is reported beside the headline and never folded into it. The headline is GCSE alone: 1,661 of 1,830 sections exact, 90.8%.
The shape of the disagreement
The headline says how often the two markers agree. It does not say what happens when they do not, and that is the part that decides whether the marking is usable. A tool that is wrong by one mark in the generous direction is a different proposition from one that is wrong by four in either. The full distribution is published rather than summarised, because summarising it is how a systematic tendency gets hidden. This table and the two paragraphs under it count all three cohorts, including AS: 2,550 sections on 78 scripts, not the GCSE headline above.
| Marky's mark, against the teacher's | Sections | Share of 2,550, all three cohorts including AS |
|---|---|---|
| Three or more marks lower | 2 | 0.08% |
| Two marks lower | 3 | 0.12% |
| One mark lower | 55 | 2.16% |
| Identical | 2,314 | 90.75% |
| One mark higher | 161 | 6.31% |
| Two marks higher | 13 | 0.51% |
| Three or more marks higher | 2 | 0.08% |
Marky is the more generous marker, and by a consistent amount
Across all 2,550 sections of the three cohorts, including AS, Marky awarded more than the teacher on 6.9% and less on 2.4%. That is a real bias and it points one way: where the two markers differ, the student is roughly three times more likely to gain a mark than to lose one. Over a whole script it comes to a mean of 1.6 marks more than the teacher gave, on papers marked out of 75 to 80; on the GCSE papers alone it is 1.5. A department that finds its Marky totals sitting a mark or two above its own is seeing the expected behaviour, not a fault. On our reading of the mark scheme, most of that generosity is earned. Of the 110 disagreements our reading settles in Marky’s favour, 99 are sections where Marky awarded the higher mark: marks the student had earned and the teacher had not credited. Marky’s own errors carry no such direction: on the 54 our reading settles for the teacher, Marky is too high on 27 and too low on 27. So, on our reading, the tendency above is mostly the marking being right rather than being soft, though it is published as a bias because a department should judge that for itself.
What that means for the review that follows
Across all three cohorts, including AS, an average script of roughly 33 sections had three come back with a mark the teacher would not have written, and the whole-script total sat a mean of 2.0 marks away from the teacher's. Four sections in all 2,550 carried a difference of three marks or more. Those are the numbers to plan review time around, and they are the reason the marking is presented as a first draft with its reasoning attached rather than as a finished result.
One marking pass, then automated checks
Every paper is marked once against the official mark scheme. It then goes through a set of automated checks, each aimed at a place where marking tends to go wrong. On maths papers, a section that may have been misread is read again from a zoomed image of the page, and what was read is corrected; the re-reading does not change the mark by itself. If a question comes back without a mark, it is asked for again, including on a stronger model; if it still has none, it is given no marks and flagged for the teacher to mark. No mark reaches a teacher with one of those checks still open.
These checks are independent of the first marking pass, not of us. They are our software checking its own output, and they are not a second opinion in the sense a moderation meeting would mean. What makes a mark checkable is not the automated pass but the teacher who reads it, which is why the reasoning is published alongside every award.
The reasoning behind each award stays attached to the question. A mark that arrives without a justification cannot be checked, and a marking system that cannot be checked cannot be trusted with a reported grade. Because the reasoning is recorded at section level, a teacher reviewing a paper can read why a mark point was given or withheld and form a view in seconds rather than re-marking the question from the beginning. The how it works page sets out the full sequence from upload to returned results.
Targeted review, with the reason stated
Every flag carries a stated reason, and none of them is the marking model's own sense of how confident it felt. In practice almost all of them are the first one below. The reason is shown alongside the paper, so the review starts from a stated concern rather than from scratch.
Near a grade boundary
A total sitting close to a boundary is where a single mark changes the reported grade, so the mark carries more consequence than its size suggests.
Possible integrity concern
Raised for a teacher to consider. A flag of this kind is advisory and does not alter the mark on its own.
A section left unmarked
A section the system could not mark is flagged to the teacher as needing a mark, and the paper is flagged so it cannot be signed off unnoticed.
Most papers sit near a boundary, so the order matters more than the flag
Grade boundaries sit roughly ten percentage points apart. So a rule that flags any total within five points of one will catch most of a class set. On our own measurement the average GCSE paper differs from the teacher by 2.0 marks. On a 75-mark paper that is about three points, so a large share of any real cohort genuinely does sit within one marking error of the grade changing. That is a fact about the grading scale, not a finding about the software. It means a flag on its own tells a department very little. So the queue is ordered rather than filtered. Papers are ranked by how close the total sits to the nearest boundary, closest first, and a department works down that list in the time it has. The papers where a single mark moves the reported grade come at the top. A few rarer reasons sort above proximity, and those genuinely are a small subset, such as a possible integrity concern or a section the system could not mark.
Who is right when they disagree
On our reading of the mark scheme, Marky is the closer of the two markers. Read that way, it is exact on 96.1% of these GCSE sections; the class teacher is exact on 93.1%. The headline figure claims none of that credit: it counts every disagreement against Marky, including the 110 our reading settles in Marky’s favour, which is why 90.8% leans low rather than high. It still moves by about half a point between runs of the same papers. Where the two differ, our reading of the scheme sides with Marky on 110 sections and with the teacher on 54. Here is that comparison with its working shown.
We read the 1,830 GCSE sections against the published mark scheme, in an answer key built in July 2026. We are the adjudicator in the last two columns, which is why the unadjudicated column sits beside them. The reading was done by AI models working for us, a different model from the one Marky marks with, not by an examiner or a teacher. The reading started from the teacher’s own mark and departs from it only where the scheme says otherwise. Most of those departures raised the teacher’s mark, the same direction Marky leans, so treat these two columns as our reading, not proof. What makes it checkable is that every disagreement names the rule it turns on. One of them reads: the student answered 77.6; the mark scheme accepts any answer between 75 and 81, and awards two marks for it; the teacher gave one. Anyone can look that up in a document neither we nor the teacher wrote.
| Cohort | Sections | Marky matches the teacher no adjudicator | Marky matches the mark scheme our reading | The teacher matches the mark scheme our reading |
|---|---|---|---|---|
| GCSE mathematics, Paper 1 (Higher) | 900 | 90.8% | 96.1% | 93.1% |
| GCSE mathematics, Paper 2 (Higher) | 930 | 90.8% | 96.1% | 93.0% |
| Both GCSE papers | 1,830 | 90.8% | 96.1% | 93.1% |
Neither marker matches the scheme on 17 sections: 5 of them are disagreements, and on the other 12 the two markers agree with each other and our reading says both are wrong. On our reading, the gap is 3.0 points in Marky’s favour.
Both are published because neither is sufficient alone. The 90.8% in that table is what we can prove without asking anyone to trust our judgement, and it is the one to hold us to. The 96.1% is what we believe the mark scheme supports, and it depends on our reading being right. A department that wants to settle it should do what we would do: take one class set, mark it, and compare question by question. That is a few hundred marking decisions, and it is worth more than either figure here.
A second teacher, on 75 more papers
Everything above rests on one teacher. So we did it again with another. 75 GCSE scripts, marked in red on the paper by a different teacher, with their own total written on the cover. We compared that total to Marky’s on every script. The two sets of papers are marked out of different totals, so the gap is also given per 100 marks, which is the only way to put the two columns side by side.
| First teacher | Second teacher | |
|---|---|---|
| Papers | 60 | 75 |
| Paper marked out of | 75 | 80 |
| Marky’s total, against the teacher’s | 1.5 marks higher | 2.8 marks higher |
| The same gap, per 100 marks | 2.0% | 3.5% |
| Marky higher on | 40 papers | 63 papers |
| Marky lower on | 10 papers | 9 papers |
| Compared question by question | Yes, 1,830 sections | No. This teacher marked in blocks |
| Also checked against the mark scheme our reading | Yes, 96.1% of sections | Yes, 225 of the 237 disagreements |
The same mild generosity, against two different teachers. One teacher marking tightly would explain one column. It does not explain both.
The question-by-question row is why the first set stays the headline. This teacher marked in blocks of one to three questions, so the second set is compared block by block rather than section by section, and the two cannot be added together. What it adds is size and a second opinion: 135 papers now carry marks written by someone other than us, against 60 before.
The disagreements, read against the mark scheme
Marking in blocks gives 737 of them across those papers. Marky and the teacher already agree on 500. Of the 237 where they differ we read 225 against the published Edexcel mark scheme, by hand, one question at a time. We are the adjudicator here, which is why it is said in the same breath as the number.
| The mark scheme sides with Marky | 161 blocks, 72% |
|---|---|
| The mark scheme sides with the teacher | 38 blocks, 17% |
| The mark scheme sides with neither | 26 blocks, 12% |
The ratio is not the part that matters. Marky sits above the mark scheme on 20 of these blocks. The teacher sits above it on 20. Exactly the same number. The whole of the difference is on the other side: the teacher is below the scheme on 167 blocks, and Marky on 44.
That is not a soft marker set against a strict one. It is a marker that credits the working, set against one who more often credited only the answer. One question on Paper 1 shows the shape of it. Five students worked out that 150 of the 400 jars were empty, split those 150 in the ratio 3 : 4 : 8, and found that 30 of the empty jars were small. Every one of the five then divided 30 by 150 instead of by 400. The mark scheme awards three of its five marks for the three steps they got right.
Totalled up, the teacher is 3.2 marks a paper below the mark scheme on these blocks and Marky is 0.5 below it. Marky is within one mark of the scheme on 95% of them; the teacher on 72%. The pattern is the same on all three papers, with the scheme siding with Marky on 69%, 72% and 74% of the disagreements on Papers 1, 2 and 3, so no single paper is carrying it.
Two things worth knowing, given here rather than in a footnote. These are blocks of one to three questions, not the question sections used elsewhere on this page, so the percentages above cannot be added to or compared with the 96.1%. And only the disagreements were read. The 500 blocks where both markers already agreed were not checked, so this says which marker the scheme prefers where the two differ, not how often either is right overall. Four scripts answered in blue ink are left out: UK exam rules ask for black ink, with pencil for diagrams, and the marking removes everything else, so neither marker was looking at the same page.
AS mathematics, read against the mark scheme
The AS cohort now has a reading of the mark scheme of its own. We marked its 18 scripts again and read every section where Marky and the teacher disagree against the published AQA mark scheme: 64 of them, each one read without sight of either marker’s mark. We are the adjudicator here too.
| The mark scheme sides with Marky | 39 of the 64 |
|---|---|
| The mark scheme sides with the teacher | 21 of the 64 |
| The mark scheme sides with neither | 4 of the 64 |
Across the whole cohort, counting the 656 sections where the two markers agree as right by the scheme, Marky matches the mark scheme on 96.5% of the 720 AS sections and the teacher on 94.0%. That is the GCSE result again, on a different exam board and a different level.
This comparison uses the latest marking run, on which Marky and the teacher give the same mark on 656 of the 720 sections. The AS row in the cohort table above uses the run before it, which matches on 653. Three sections either way is the ordinary variation between two runs of the same papers.
Physics, read against the mark scheme
Physics has no full set of teacher's marks to compare against, so we went further than reading the disagreements. We read every section of 15 AQA physics papers against the published mark scheme, 9 at A-level and 6 at AS: 605 question sections in all, each one read without sight of Marky’s mark. We are the adjudicator here. The reading was done by AI models working for us, a different model from the one Marky marks with, not by an examiner or a teacher.
| Marky matches the mark scheme | 564 of 605, 93.2% |
|---|---|
| Within one mark | 600 of 605, 99.2% |
| Multiple choice | 230 of 235, 97.9% |
| Written answers | 334 of 370, 90.3% |
Marky matches the mark scheme on 93.2% of the physics sections and is within one mark on 99.2%. Over a whole paper it sits 0.7 marks above our reading, on papers marked out of 35 to 85 (one A-level Paper 3 was marked in its two parts, which count separately here).
This is a stricter test than the 96.1% above, and the two should not be set side by side. That figure reads only the sections where two markers disagree and counts the rest as right. Here every section was read, and nothing was assumed.
Every accuracy figure here is measured against a teacher, or labelled as our reading
The first mark-scheme table covers the 1,830 GCSE sections. The AS cohort has the teacher's own marks, and they are what the AS row of the cohort table above uses, but they were supplied on a spreadsheet rather than written on the scripts. The AS scripts themselves are unmarked: the cover's examiner grid is blank and there is no pen on the pages. The GCSE scripts are the opposite, marked in red with a total on the cover. The AS answer key we held before was built by starting from Marky's marks and editing them, so 96% of it was still Marky's own output, and scoring Marky against it would have been scoring Marky against itself. That is why the AS and physics readings above were made from scratch, without sight of Marky's marks. We have marked a good deal more than appears on this page: other GCSE papers, other subjects, and cohorts going further back. No accuracy figure from any of it is published here. The references we hold for most of those sets were built by starting from Marky's own marks, so a percentage scored against them would be partly Marky agreeing with itself, and the rest cover too few papers to carry a figure. They are useful for comparing one version of the software with another, and that is all we use them for. Every accuracy figure on this page is measured against a mark a teacher wrote, apart from the mark-scheme figures, which are our reading.
Neither figure is a claim that a teacher marking at ordinary speed should reach the mark scheme every time. Marking a class set is done in the evening against a deadline, which is the entire reason this product exists.
The teacher holds the final mark
Any mark can be changed by the teacher before results are published, on any question, with or without a flag. Nothing is locked, and no result reaches a student until the teacher releases it. The marking is a first draft produced quickly and consistently, and professional judgement remains where it belongs.
That is what makes the figures above the conservative reading. They describe the marking a teacher receives, before the teacher has looked at it. What a class actually receives has been through a professional who can change any mark, so we expect the released result to be more accurate than any number on this page. We have not measured that, so it is an expectation, not a figure. A department that overrides a handful of marks per class set is using the system exactly as intended.
What the figures cover
Every figure above comes with its scope, so a department knows exactly what it is getting.
Measured on mathematics and physics
The 1,830 headline sections are GCSE mathematics, and the 720 AS sections reported beside them are mathematics too, all measured against the teachers' own marks. Physics is measured separately, against our reading of the mark scheme, on 605 sections at AS and A-level. No accuracy figure is published for any other subject, and nothing on this page should be read across to one.
Measured against a real class teacher
The marks compared against are what a teacher wrote on their own class set, working at ordinary speed. That is the right reference for a department deciding whether to use this, because it is the marking Marky would be replacing. It is not a standardised examiner judgement, and it is not error-free. Some of the sections counted as wrong are places where Marky was right and the teacher slipped. On the 1,830 GCSE sections the two markers disagree on 169, and read against the published mark scheme by us, it sides with Marky on 110 of them and with the teacher on 54. The headline does not subtract them, because subtracting them needs our own adjudication, and that is published in its own labelled columns rather than folded into the number. The GCSE teacher marks were also transcribed from the ink on the script rather than supplied as a data file, so a transcription slip shows up here as a disagreement, an error that counts against Marky, not for it.
The headline rests on nothing we produced
The strongest thing that can be said against any AI marking claim is that the company decided what the right answer was. The headline, 90.8% on GCSE alone, rests on nothing we produced. The AS marks were written with Marky's mark in view, which is why AS is reported beside the headline and never folded into it. Where we do adjudicate, in the mark-scheme columns, it is labelled as ours in the same sentence as the number and it sits beside the unadjudicated figure for the same sections. Physics is the one exception: there is no full set of teacher's marks to compare it with, so its reference is our reading of the mark scheme, and it is labelled that way wherever it appears.
Every mark is shown with its working
Marking begins with reading, and a paper that a human marker would struggle over is a paper the system will struggle over as well. Faint pencil, heavy crossing out, working that runs across two pages and a photograph taken at an angle all make a section harder to read correctly. The effect on accuracy is real and varies with the class and the scan, which is why every mark is shown with the working it was read from.
How the AS marks were collected
The AS scripts came to us unmarked. The teacher's marks were collected afterwards, on a spreadsheet with one row per question and four columns: the student, the question, the mark Marky had given, and an empty column for the teacher's own mark. So the teacher could see Marky's answer while writing theirs. We checked that they were marking rather than copying, and they were. Two things show it. They disagreed with the mark in front of them on 63 of the 720 sections, 67 against the run the AS row above uses. And their marks added up to a separate list of paper totals on all 18 scripts. Neither check rules out a nudge toward the mark they could see. The 1,830 GCSE sections carry no such exposure. That is why the headline is GCSE alone, 90.8%, and AS is reported on its own.
What the 78 scripts are
Seventy-eight scripts across two qualifications, in one subject: 30 GCSE Paper 1, 30 GCSE Paper 2 and 18 AS. The headline uses the 60 GCSE scripts; the 18 AS are reported separately. The two GCSE papers came from the same year group, so many of the same students appear in both. They are set out exactly so a department can judge how close this is to its own class sets. The one that settles it is ten papers of your own, which is what the trial is for.
What the physics figure is
Fifteen AQA physics papers, 9 at A-level and 6 at AS: 605 question sections. Some scripts carry a teacher's ink, but it was never collected into a full set of marks, so the reference is our own reading of the published mark scheme, made section by section without sight of Marky's mark. That makes it a check against the scheme rather than against a teacher, and because the reading is ours it is not an independent one. On it, Marky matches the scheme on 93.2% of sections and is within one mark on 99.2%. It is reported on its own and never folded into the mathematics figures.
Every figure on this page is one we would rather defend than inflate: a number that survives being checked is worth more to a department than a bigger one that does not. The same reasoning appears in the comparison with manual marking, where the honest baseline is not perfect marking but tired marking on a Sunday evening.
What is marked, and how quickly
A class set of biology papers takes approximately five hours to mark by hand. On Instant most classes come back in 20 to 30 minutes, and on the Standard lane it typically takes 60 minutes.
Marking covers mathematics, biology, chemistry and physics at GCSE, AS and A-level, against official mark schemes from AQA, Pearson Edexcel, OCR and WJEC/Eduqas. Papers are accepted as PDFs and as photographs in .jpg, .png, .heic or .webp with no conversion required, so a set can be captured on a phone at the end of a lesson.
The trade is speed against allowance, and the arithmetic is worth stating plainly. Both lanes are charged at the same £1.20 per paper drawn, but Instant draws two papers per script where the Standard lane draws one. A script marked on Instant therefore costs £2.40 rather than £1.20, and a class set of 30 uses 60 papers from the allowance rather than 30. Instant also typically returns a single paper on its own in about 15 minutes.
Instant is available once it is switched on for the school, and is then chosen for each marking run. Both timings describe typical runs rather than a guaranteed turnaround, and marking can take longer at busy periods. The annual department allowance starts at 1,200 papers.
Full details of allowances and top-ups are set out on the pricing page, and the marking features that sit on top of the results are listed under features.
Questions departments ask about accuracy
What counts as an exact match?
An exact match means the mark Marky awards for a single question section is identical to the mark the class teacher wrote on that section. It is measured per section rather than per paper so that compensating errors cannot cancel out.
How large is the comparison, and who marked the reference?
The headline is 1,830 question sections on 60 GCSE scripts: mathematics Paper 1 (Higher) with 900 sections and Paper 2 (Higher) with 930, marked by the teacher in red pen before Marky saw them. A further AS mathematics cohort, 720 sections on 18 scripts, is reported separately, because that teacher's marks were collected afterwards on a spreadsheet that showed Marky's mark. The reference marks are the teachers' own throughout. No examiner was involved, and nobody adjudicated between the teacher and the software.
Does every paper have to be checked by a teacher?
No. The set comes back ranked by how close each total sits to a grade boundary, with the reason stated, so a department reviews in priority order rather than paper by paper. The order is by how much one mark would matter, not by how likely the marking is to be wrong, so it is worth spot-checking papers from across the list too. Any paper can be opened at will, and nothing reaches a student until the teacher releases it.
Can a teacher change a mark?
Yes. Any mark on any question can be overridden before results are published, and nothing reaches a student until the teacher releases it.
Which subjects were measured?
Mathematics at GCSE and AS, against the teachers' own marks, and physics at AS and A-level, against our reading of the mark scheme. No accuracy figure is published for any other subject Marky marks.
Does Marky give the same grade as the teacher?
We do not publish a figure for that, because the two GCSE cohorts are a school's own end-of-year papers with no official grade boundaries, and any grade agreement number would rest on boundaries we had invented. The whole-script total is as far as the evidence goes: on the GCSE papers, a mean of 2.0 marks from the teacher's total, on papers marked out of 75.
The benchmark that counts is the department's own
Every figure on this page comes from someone else's class set. The one that will decide it is yours. During the free 14-day trial, submit ten papers the department has already marked and compare them question by question, on the specification you teach and the handwriting you actually get. These figures are published because we expect them to hold on your papers, and a familiar class set is the fastest way to find out. No card is required.
The blog sets out a five-step method for running that comparison, including how to avoid the trap of comparing totals and the three kinds of disagreement to expect.