Skip to main content

How accurate is our IELTS scoring? Same score, every time.

In the run of 19 September 2026, BandNine scored 24 IELTS essays with a mean absolute error of 0.33 of a band against our own descriptor-based labels. No script in that set has been marked by a certified IELTS examiner, so this measures agreement with our labelling, not with the examiner who will mark your test. We re-run it every Saturday and publish the result here, good or bad.

The consistency contract · test it yourself

Submit the same essay twice. You'll get the identical result, word for word.

The first time an essay is scored, that result is sealed as canonical for that essay and task. Re-submit it a minute later or a month later and the sealed result is returned byte-for-byte: same band, same criterion scores, same feedback, same corrections. Changing spacing, capitalisation, or curly quotes doesn't fool it; the essay is content-addressed, not text-matched.

Three honest footnotes: the quick preview that streams in first is provisional. The final score that replaces it seconds later is the sealed one. A different essay (even one changed word) is scored fresh, because one word can legitimately move a band. And the seal is keyed on the scoring model version as well as the essay and the task, so when we upgrade the model your essay is re-scored instead of replaying an answer from the old one. That is deliberate: a seal that outlived every upgrade would freeze you on whichever model happened to mark you first. If you ever see the same essay produce two different final scores, that is a bug and we want to know: hi@bandnine.ai.

How accurate is it this week?

Every Saturday we re-score our in-house, descriptor-labelled calibration set through the production scorer and publish the deviation. We have not found another AI IELTS scorer that publishes one. We do because at this price, you deserve to see the numbers.

Latest measurement · 19 September 2026
Mean Absolute Error
0.33
Within half a band
bands deviation vs our calibration set, labelled criterion by criterion against the public band descriptors
00.5 · within half a band1.0 · close1.5+
This is within half a band of our calibration reference. We do not claim examiner-level accuracy: that would need scripts marked by certified examiners, and we do not have those yet.
Task Response0.52
Coherence & Cohesion0.33
Lexical Resource0.33
Grammar Range & Accuracy0.33
Worst single essay0.50 bands
Calibration set size24 essays
Drift vs prior week-0.05
Green · MAE ≤ 0.5

Within half a band of the calibration reference. Tight enough to guide practice with confidence.

Amber · 0.5 to 1.0

Within one band. Reliable for guiding practice, not yet exam-exact.

Red · above 1.0

Off by more than a full band. We publish it and are actively tuning the model. No hiding the miss.

The full accuracy record

We publish our AI's scoring accuracy every week. Here's the data.

Every run since May, the method in full, how the calibration set is graded, and an honest list of what this number does not promise.

Read the data

What stops a bad score reaching you?

Six mechanisms, each one live in production code, that keep the scoring honest, consistent with the IELTS rubric, and free of hallucinated feedback.

Sealed first score

Every (essay, task) pair gets one canonical result, stored against a content hash with the model version pinned and temperature locked to 0. Re-submissions replay it exactly, so scores can't wander between attempts.

Official IELTS rubric, enforced

The examiner prompt embeds the public IELTS band descriptors for all four criteria. The maths is then re-checked server-side: every band must sit on the official 0.5-step scale, and the overall must equal the rounded mean of the four criteria. If the model returns anything else, the server corrects it before you see it.

No invented quotes

When feedback claims you wrote something, the claim is checked verbatim against your actual essay or transcript. Any "correction" quoting words you never wrote is dropped before the response leaves the server.

Deterministic error scan

A rule-based grammar layer (not AI) scans every essay and appends high-confidence errors the model under-reported. Rules are pure functions: the same essay always produces the same hits.

Drift alarms + a build-blocking guard

If a live re-score of a known essay ever drifts, it's logged and alarmed. And every deploy runs an automated determinism check. If any of these guarantees regresses in code, the build fails before it can reach you.

Weekly benchmark against a fixed set

The accuracy report above re-scores the whole calibration set through the exact production pipeline every Saturday and publishes the deviation, including the weeks we don't like the number. Rows are never edited or deleted.

How is this measured?

We maintain a calibration set covering Task 1 and Task 2, Academic and General Training, from band 4 upward. Each script carries a per-criterion ground-truth band (Task Response, Coherence & Cohesion, Lexical Resource, Grammatical Range & Accuracy) plus written notes quoting the descriptor behind it. That ground truth is currently in-house, descriptor-labelled, which is the single biggest limitation on the number above and is spelled out in full in the accuracy report.

Every Saturday at 06:00 UTC, an automated job re-scores every script in the set through the production scoring endpoint, the same code path real students hit, with the same task question attached and the sealed-score cache bypassed so nothing is replayed. It computes the mean absolute error per criterion, plus the worst single-script deviation, and stores the result with the full per-script detail in a public table.

We now publish the direction of the error as well as its size. Mean absolute error cannot tell "consistently half a band too generous" from "half a band either way at random", and those are not the same product. Where a criterion leans by more than a quarter of a band across the set, the panel above says so.

Thresholds, and what happens when one is crossed. Overall MAE above 0.5 bands (one step on the IELTS half-band scale), any single script more than 1.0 band out, or a criterion leaning more than 0.35 bands in one direction: the build fails and the drift has to be fixed before any further change ships. Previously this only opened a ticket, and three such tickets stayed open from June to August, so the check now blocks rather than notifies. A run that scores less than four fifths of the set is treated as an infrastructure failure and refuses to publish at all.

One more thing you are entitled to know: a script can be quarantined out of the set if its ground-truth label turns out not to be defensible, and that changes the figure above. One script is quarantined today (an Academic Task 1 report labelled Band 6.5 whose own notes describe Band 7.5 to 8 work). Every quarantine carries a written reason in the same public table, and the build fails if one does not.

We have not found another AI IELTS scorer that publishes this. We do because verifiable accuracy at this price is the actual differentiator.

How accurate are AI IELTS band scores?

The questions people ask before trusting any of this, answered on the page that holds the evidence rather than on a sales page.

How accurate are AI IELTS band scores?+

Ours is currently 0.33 of a band of mean absolute error, measured over 24 reference scripts whose bands were labelled in-house, descriptor-labelled. That figure is re-run every Saturday and published here whether it improves or not. Treat any AI scorer that cannot tell you its sample size and who marked the reference set as unmeasured rather than accurate.

Can you trust an AI IELTS writing checker?+

Trust the ones that show their working. Three questions separate a measurement from a marketing line: how many scripts is the figure measured over, who marked those scripts, and what does the figure not cover. We answer all three on this page, including the part that costs us something, which is that no certified IELTS examiner has marked any script in our set.

Is an AI IELTS band score accurate enough to plan my test date around?+

Use it to see whether your writing is moving, not as a prediction of the band you will be awarded. Two things make the official result different: it is marked by certificated examiners, and the exam board publishes no error figure for Writing, so no tool can honestly state its distance from an official band. What a scorer can honestly show you is consistency, a published reference set, and whether your own answers are improving against it.

Which IELTS writing checker is the most accurate?+

Nobody can answer that yet, and we are not going to pretend we can. It would need every tool measured against the same scripts marked by the same people, and no such comparison exists publicly. What can be compared today is disclosure: what each tool publishes about how its own accuracy was measured. Several publish a percentage. We could not find one that publishes the sample size behind it.

What we do not claim

A measurement is only worth reading if you know what it excludes.

  • That our scoring matches a certified IELTS examiner. No examiner has marked any script in the reference set.
  • A percentage accuracy figure. On a band scale a percentage does not say accurate to within what, over how many scripts, or marked by whom.
  • That we are more accurate than any named competitor. That needs both tools measured on the same scripts by the same markers, and no such test exists.
  • That the figure covers every part of the exam equally. General Training Task 1 holds too few reference scripts to support a claim, and we say so rather than averaging it away.
  • That the number will not get worse. It has, publicly, and the week it hit 1.20 is still in the table below.

What has the number done over time?

DateEssaysMAEStatus
19 Sept 2026240.33 within half a band
12 Sept 2026240.38 within half a band
5 Sept 2026240.35 within half a band
31 Aug 2026240.31 within half a band
29 Aug 2026240.44 within half a band
23 Aug 2026240.33 within half a band
22 Aug 2026270.44 within half a band
15 Aug 2026270.39 within half a band
8 Aug 2026270.37 within half a band
1 Aug 2026270.41 within half a band
25 Jul 2026280.39 within half a band
18 Jul 2026280.41 within half a band

Want to verify? Every figure on this page, plus the per-script detail behind it, is served without a key at /api/quality/calibration?history=1. The underlying tables are public-read in Supabase by design. Are you a certified IELTS examiner willing to mark scripts for the set? That is the contribution we most need. Reach out at hi@bandnine.ai and we’ll send the schema.

Read the full accuracy report →Try the scorer yourself →