The accuracy figure we publish every Saturday covers Writing. It does not cover Speaking. Stretching that number across to a spoken answer is the easiest sentence here to write and the least defensible, so it will not appear. What follows is what is actually known, and what a practice speaking band is worth.
#How accurate are AI IELTS speaking scores?
Nobody can give you an honest single figure, and that includes us: our published weekly accuracy figure covers Writing only, and we publish none for Speaking, Listening or Reading. This week's figure sits on our quality page. It is the mean absolute error between the band our scorer gives a script and the band on our label for it, recomputed automatically every Saturday at 06:00 UTC through the live production scoring endpoint with the sealed score cache bypassed.
Two things matter before anyone quotes it. The calibration set holds 24 active scripts, labelled in house by the two of us against the public IELTS band descriptors. Not one has been marked by a certified IELTS examiner: the endpoint returns a mix of zero examiner and 24 curated, and the raw JSON is public at bandnine.ai/api/quality/calibration. Our internal pass mark is a mean absolute error of 0.50 of a band, and a worse run, or any script more than a full band out, opens a ticket automatically. So when a tool quotes you a percentage, ask what it was measured against.
#Why is speaking harder to score automatically than writing?
Speaking is harder because the software never hears your answer the way a person does: it scores a transcript, so any word the transcriber gets wrong becomes a vocabulary or grammar error you never made. Writing arrives as text. Speech must be converted first, and that conversion is a second place for things to go wrong.
The failure is specific. You say "I'd been living there for two years". The transcriber writes "I been living there for two years". The scorer, reading only text, marks grammatical accuracy down for an error that existed nowhere except the transcript. Say "refurbished", have it heard as "furnished", and your lexical credit disappears. False starts, self corrections and filled pauses flatten into a wall of errors, where a trained listener hears a false start as a false start.
IELTS says much the same. Its insight article "What automarking means for language test validity and integrity", published 13 May 2026, states: "Generally, automarking is more achievable for writing tasks than speaking." Note what that article is: guidance on what to weigh up "before introducing a new scoring system". It never says IELTS uses automarking on the live test.
#Which speaking criteria can an AI actually judge?
A tool that measures the audio can read pace and hesitation well. A tool that scores the transcript alone, ours included, sees only what survives into text, which is why pronunciation is the criterion it reads worst of all. Knowing which is which tells you how much weight to give each part of your feedback.
| Criterion | Seen reliably | Where it breaks |
|---|---|---|
| Fluency and coherence | Repetition, self correction and filler words, wherever they reach the transcript. Pause length and speech rate need the audio, so a scorer reading only text does not see them. | Coherence. Whether your ideas develop is a higher order judgement. |
| Lexical resource | Range, repetition, obvious collocation errors, reaching past common words. | Precision in context. Less common words are the ones a transcriber mangles. |
| Grammatical range and accuracy | Structures you attempt, and errors that repeat across sessions. | Most damaged by transcription. Dropped auxiliaries and endings are the classic false positive. |
| Pronunciation | Very little, if the tool scores from text. | Everything the criterion is about. A transcript cannot show stress, chunking or intonation. |
IELTS names the same weakness: the limitation is "the ability to evaluate higher-order skills such as organisation, idea development and the nuances of human communication". IDP is blunter: "Some AI tools may focus on grammar or pronunciation, but miss nuances of clarity, style or coherence." So comments that quote your own words back to you, and errors that show up in recording after recording, are the sturdy part of your feedback. Comments on how well your argument hangs together are the part to hold loosely.
#Will my accent lower my AI speaking score?
It should not, and by the descriptor it must not: the IELTS pronunciation criterion is built around how easily a listener understands you, not around whether you sound British, Australian or American, so a scorer that marks down a clear regional accent is wrong by the rules it claims to apply. An accent a listener follows without strain is not a pronunciation problem.
That is the rule, and the rule does not enforce itself. Automatic transcription is not equally accurate for every speaker, and where it is worse the scorer reads a worse answer than the one you gave. We have not measured which speakers that hits hardest, so we will not put a cause or a number on it. IELTS states the consequence: "The inconsistency in automarking accuracy across test takers would lead to bias in test scores if the automarker is the only means" of scoring. Hence its stated position, a hybrid with human oversight, citing the UK regulator Ofqual's rule that AI "cannot be the only means of determining results for high-stakes qualifications".
What we do about it is mostly what we refuse to do. We publish no accuracy figure for Speaking, and none per accent, because we have not measured either. What we can offer is a habit: if a tool shows you a transcript, read it before the band, and treat any comment attached to a word you did not say as void. A tool that never shows you what it thought it heard cannot be audited.
The real test is not scored this way. IDP states that "IELTS Speaking is scored live with a real, human examiner", and IELTS.org says every test is assessed by qualified examiners with rigorous training and continuous monitoring, with selected results marked twice.
#Can AI identify a memorised answer?
Sometimes, but not reliably, and never the way a person in the room can. Software can flag the pattern many memorised answers share: unusually fluent delivery, a register that jumps between parts of the test, content that answers a slightly different question from the one asked. Those are inferences from a transcript, not detection. A memorised answer that fits the question and is delivered naturally reads like a strong one.
A live examiner has two things a transcript does not. The first is delivery: recited language has a flat, running on rhythm, and it does not react to the person opposite. The second is that Part 3 is a discussion, so the follow up depends on the sentence you just said, and a prepared answer meets a question written one second ago. That is our read as teachers rather than a published finding.
Memorising whole answers is a bad trade even where nothing catches it, because a fixed answer rarely fits the question, and steering it damages the coherence you were protecting. Memorise flexible chunks instead: language for buying thinking time, comparing, speculating, correcting yourself cleanly. Then practise them on unfamiliar prompts, using our speaking cue cards and Part 2 strategy guide.
#What is a practice speaking band actually good for?
A practice speaking band is a direction indicator, not a prediction of your result. It can tell you which of the four criteria is dragging the others down, whether that criterion moves over weeks, and whether the same grammar error turns up in every recording. It cannot tell you whether to book or postpone your test, predict your result to the half band, or settle an argument about a result you already have.
Treat one practice speaking band as noise and three as a signal: if the same criterion comes out lowest in three sessions running, that is your real weakness, and if a band jumps once and then comes back, it never moved.
Track the criterion, not the overall band. An average of four numbers hides the movement you are working for.
#What do the test owners and the research actually say?
Both IELTS co-owners describe Writing and Speaking as marked by trained human examiners, and the study everyone reaches for is a Writing study, not a Speaking one. IDP tells candidates to "Refrain from using LLM tools to grade your IELTS essays", because the result "will not match up to how human examiners will mark your Writing test, according to official band descriptors". It warns that AI "is well-known for justifying its mistakes and giving overwhelmingly positive feedback", and advises: "Focus on using AI for correction and feedback, but not creation." The IELTS automarking article carries the line that best frames any practice tool, ours included: "Using automarkers in low-stakes practice contexts has very different implications from a high-stakes university entry test."
The study people quote is Koraishi (2024), in Language Teaching Research Quarterly, volume 43, pages 22 to 42. It compared ChatGPT 4, the November 2023 version, with published official band scores on 55 real Writing Task 2 essays, reporting an intraclass correlation coefficient of 0.814, 95% confidence interval 0.702 to 0.887, and a weighted kappa of 0.811. Notice what it is not: 55 essays, Task 2 only, no Speaking tested, and the author's own conclusion that "ChatGPT should not be implemented as an official rater, at least not yet". Citing it as proof that AI can score your speaking extrapolates past both the module and the finding.
#One thing we can promise
The first full score for a piece of text is sealed, and every later submission of that same text returns it byte for byte: same band, same criterion scores, same feedback, enforced by a check in our build pipeline. This is a consistency property, not an accuracy one, and a scorer can be consistently wrong. It does mean a band cannot be farmed by resubmitting: when your score moves, something in your answer moved.
The free tier includes three AI speaking scores and three writing scores, with feedback visible but blurred. Record answers to unfamiliar prompts rather than rehearsed ones, and keep three sessions before drawing a conclusion. If pronunciation is your lowest criterion, read the pronunciation guide. For the Writing side, see how accurate our AI IELTS scoring is and whether ChatGPT can score an essay.