We spent a while reading the rest of the market properly: the specialist marking platforms, the general AI toolkits teachers already have logins for, the university assessment systems, the essay graders that plug into a classroom LMS (learning management system, so Google Classroom, Microsoft Teams or Satchel One). It is a genuinely crowded field now, and there are some good products in it.
It is also a field where every product page makes the same six promises. Marks to your mark scheme. Reads handwriting. Saves you hours. Aligned to AQA, Edexcel and OCR. Teacher stays in control. Detailed feedback in minutes. Read four of them in a row and they blur.
The differences are real, they are large, and almost none of them are visible on the homepage. Below are the questions that find them. This is our blog, so we answer each one for Marky too, including the places where the honest answer is a limit rather than a win.
1. What does it actually take in: a whole script, or an answer?
This is the first fork in the road and it eliminates half the market for most departments.
A good number of tools are built around one piece of work: an essay, typed into a box or pasted from a document. Several are built around a photograph of a single page. The LMS-connected graders assume the work arrived as a digital submission in the first place, because that is what an LMS holds. These are real products solving real problems, and for a Year 10 English essay set as homework they can be excellent.
None of that is the mock your class sat in the hall. That is a stapled twelve-page booklet, front cover to back, with working in the margins, a crossing out on page four, a graph on page seven and an answer to 6(b) that continues on the spare page at the end.
Ask for: upload a real class set of scanned booklets during the trial. Not the vendor’s demo paper. Yours.
Marky: the whole script, scanned on the department photocopier or photographed on a phone, PDFs or images, no conversion step. The full booklet is read straight from the scan, working included, and marked question by question.
2. Whose mark scheme is it actually using?
“Marks to the mark scheme” covers three quite different things.
- A rubric you paste in. The tool sends your criteria to a general purpose language model and asks it to apply them. Quality tracks how well the model reads your rubric, and it does not know a mark scheme convention from a bullet point.
- A model calibrated to board material. The vendor has tuned or benchmarked against published standardisation material for specific papers. This can work well on the papers they prepared for, and it tells you much less about the department assessment you wrote yourself last Tuesday.
- The actual mark scheme for the actual paper, supplied by you. The published scheme for that paper and series, applied point by point.
The third is the only one that survives contact with a real department, because departments do not only sit past papers. They sit past papers, edited past papers, in-house assessments and end of year exams assembled from three sources.
Ask for: does it award against a scheme I upload, and can it tell me which criterion earned each mark?
Marky: you upload the mark scheme first, then the class set. Every question gets its own mark and feedback, quoting the part of the answer it rests on and naming the marks awarded (M1, A1 and so on) where the scheme uses them, which is what makes moderating it quick rather than a re-mark.
3. Does it hold up where the answer is working rather than prose?
Look closely at subject coverage lists. They are usually longer than they are deep, and they lean heavily towards the essay subjects: English, history, geography, business, psychology, sociology, RS. There is a reason for that. Prose against level descriptors is the tractable end of this problem.
Maths and the sciences are the hard end. A mark scheme there is a machine: method marks that pay out even when the arithmetic goes wrong, error carried forward, alternative valid methods the scheme never anticipated, a sketch that earns a mark for its shape, units, significant figures, a final answer that is right for entirely the wrong reason. Getting that wrong does not look like a bad essay grade. It looks like a student losing three marks they earned.
Ask for: run a maths or physics script with genuinely messy working through the trial and read the awards, not the total.
Marky: maths and the sciences are what we built for. Maths is measured against teachers’ own marks and physics against our reading of the mark scheme; biology and chemistry are not yet measured. Scanned or typed, read straight from the page, marked the same way.
4. Who marks first?
There are three models on the market and they are sold in similar language.
- Assist while you mark. The system groups similar answers, suggests feedback as you go, or learns your marking style as you work and starts predicting it. Some of these are excellent pieces of software. Note what they are: your marking gets faster, and you still make the first pass over every script.
- Mark and publish. The AI marks and results reach students. Fast, and it puts a machine where your professional accountability is supposed to be.
- Mark first, teacher moderates. The AI does the first pass, the teacher reviews, changes whatever is wrong, and releases.
The middle option is the one to be careful with, and not only on principle. Ofqual’s position is that AI is not yet ready for high stakes marking, and the regulator and JCQ are consistent that the assessor stays accountable for the mark whatever tool assisted it. A product that pushes marks straight to students is quietly moving that accountability onto you without the review step that would justify it.
The first option is the honest one to compare against, and the arithmetic is what decides it: if you have still marked all thirty scripts, the weekend has still gone.
Marky: Marky is the first marker and you take the second marker’s chair. Nothing reaches a student until you release it. We have written about how to audit that on one class set before trusting it with anything that counts.
5. What does the accuracy number actually mean?
This is where the market is weakest, and it is the question we would push hardest on with any vendor, ourselves included.
Accuracy claims are everywhere. Correlating at over 90% with chief examiners. Half the error margin of a human marker. Calibrated against official exemplars. So we spent an afternoon trying to read the evidence behind every school facing tool we could find, and the distribution is worth reporting honestly, because it is not what we expected.
Most publish nothing measurable at all. Not a weak number or an old number: nothing. The accuracy section is a sentence saying the tool is calibrated against official mark schemes and government exemplars, or a claim that marks land within one of a senior examiner’s “in the large majority of cases”. No sample, no denominator, no definition of agreement. Several publish worked examples instead: one essay marked two or three times, landing close to itself each time, presented as evidence of accuracy. It is not. Consistency is not accuracy. A marker that runs reliably two marks light is perfectly consistent and wrong on every script.
The university platforms are a fair exception, and deserve credit. At least a couple have genuine peer reviewed research behind them. Read it and you find it measures grading time, workflow consistency and how much feedback students end up receiving. Those are real findings, honestly obtained, about a process in which a human still makes every marking decision. They are not claims about whether a machine’s mark was right, and to their credit they are not presented as such.
One vendor does publish proper statistics, with sample sizes, on named board material, including a head to head on a stated number of standardisation essays. That is the strongest published evidence we found in this market and it is more than most of the field offers. It is also built on the statistic the whole market reaches for, which is worth understanding before you weigh any of it.
A correlation is not an accuracy figure
A Pearson correlation above 0.9 with senior examiners sounds like it settles the question. It does not, and the reason matters for your students.
Correlation measures whether two sets of marks move together, not whether they agree. A marker who awards exactly two marks below the scheme on every single script correlates at a perfect 1.0 and is wrong every time. Correlation answers “does it rank the class the way an examiner would”. It cannot answer “is this 47 actually a 47”, and the second question is the one that decides a grade, a set placement and a predicted offer.
So when you are handed a correlation, ask for two more things: the mean absolute error, which is in marks and means something to a teacher, and better still the share of marks that land within one, with the denominator underneath it. A vendor who has done the work will have both. A vendor who leads with correlation because it is the flattering number will have neither.
Ask for, in writing:
- How many units, and what is a unit? Whole paper totals hide compensating errors, where a mark lost on question 3 comes back on question 7. The honest unit is the question part.
- Whose scripts? Vendor demo scripts, exam board exemplars that the model may well have been tuned on, or real class sets marked by the teacher who taught the group?
- Compared with what, and who adjudicated? If the vendor decided each disagreement, the number measures the vendor.
- What counts as agreement? Exact match, or within a mark? Both are legitimate. Quoting one and implying the other is not. If the answer is a correlation, it is not an agreement figure at all.
- Which subjects was it measured on? A figure earned on essay subjects tells you nothing about whether the method mark on question 6 survived.
- What was excluded? Papers that failed to process, questions skipped, subjects left out of the sample.
- Where is the limits section? A vendor who has measured properly knows exactly where their evidence stops. A site with no limits page has usually not looked.
Marky: our accuracy page publishes the method before the numbers. 1,830 question sections across 60 real GCSE maths scripts, compared against the mark the class teacher had already written on the page, with nobody adjudicating between the two: 99.2% of marks land within one of the teacher’s. On the GCSE cohorts, measured against our reading of the published mark scheme rather than against each other, the software agreed with that reading on 96.1% of sections and the teacher on 93.1%. The average paper differs from the teacher by 2.0 marks. Those are agreement figures on question parts, not correlations, and there is a section of stated limits at the bottom which we would rather you read first.
We could be wrong about the field. This was an afternoon of reading public pages, not an audit, and a vendor may well have a methodology we could not find. If that is you, send it to us and we will link it here.
6. What happens after the marks exist?
Marking is the part every vendor builds, because it is the part that demos well. It is not the part that eats the week.
The rest of the job is thirty-one totals typed into a spreadsheet, boundaries looked up for the right series and applied by hand, working out who needs what on Monday, writing something usable for the parent who emails on Tuesday, and having evidence ready for the next data drop. Marking a paper and handing back a marked paper leaves every one of those exactly where it was.
This is the single biggest split we found in the market. A number of the tools, including some of the sharper and cheaper ones, do the marking and stop. You get a marked script and a feedback blob, and then you are back in the spreadsheet you were trying to escape.
Ask for: show me the class set after marking, not the marking.
Marky, after the marks exist:
- The set comes back ordered by how close each total sits to a grade boundary, closest first, with the reason stated. To be precise about it: boundaries sit roughly ten percentage points apart, so most of any real class set is near one. The useful signal is the order you work down, not a flag. A section the system declined to mark rather than guess at, and a possible integrity concern where the department has asked for that check, sort above proximity.
- Grade boundaries applied automatically for the series you configure.
- Question analysis, specification coverage and misconceptions across the class, so the reteach list writes itself.
- A report per student and a gradebook export when you are ready to record marks.
- Student and parent portals, so feedback is something a student works through rather than a sheet in a bag.
- 60 languages, for the parent conversation that goes better in the family’s own language.
7. What does a class set cost, and what is the meter?
Pricing models in this market vary more than the prices do, and the model matters more than the headline number.
- Per student per year, institution wide. You pay for every student on roll whether or not a single one of their scripts is marked.
- Per teacher seat. Charges the wrong number entirely. Ten teachers who each mark two sets a year cost more than one teacher marking two hundred.
- Monthly subscription with a page allowance. The cheapest looking option on the shelf, and the one to model carefully. Two things to check. First, the meter is pages, and a mark scheme is pages: ask explicitly whether the scheme and question paper count against your allowance, because on a page meter they generally do, and a scheme can run longer than the script. Second, marking is bursty and monthly caps are not. A two hundred student cohort sitting three papers over a fortnight in November is several thousand pages inside two weeks, on a plan sized for a quiet month in October.
- Contact us for school pricing. Sometimes fine. Also the model where two similar schools pay different amounts.
Marky: £1.20 per paper excluding VAT, bought as a yearly department allowance from 1,200 papers. One paper is one student’s script marked end to end, with every verification pass included and one free re-mark of a question you query. Mark schemes and question papers are not metered. The allowance runs the full academic year with no monthly cap, so mock fortnight is a normal week. A script that fails to mark is reconciled back to nothing. There are no seats, no per teacher charge and no feature tiers: analytics, portals, exports, boundaries and languages are in it at every level, because a department that has paid for marking should not discover the reporting is an upgrade. It is bought on a purchase order against a department budget, so there is no card on file and no subscription anyone has to remember to cancel. Two honest caveats: the Instant lane, which marks a whole class set simultaneously, draws two papers per script instead of one at the same unit price, and unused papers expire at the end of the school year.
What we would do in your position
Ignore the feature grids, including ours. Take one real class set, the actual mark scheme, and the messiest handwriting in the group, and run it through two or three tools in the same afternoon.
Then compare the answers to the seven questions above, not the marked scripts alone. The marking step is where these products look alike. Everything either side of it is where they are not.
If you want to run that afternoon with us, the free trial covers ten papers with no card, from any class or subject, and the clock only starts when you first start marking.