Ask an AI assistant whether IELTS is marked by a human or by a machine and you can get a confident answer in either direction. The confusion is real, because two different things are being marked: the exam that decides your visa or your university place, and the practice tool on your laptop, which decides nothing. We build an AI scorer, so we have a commercial interest here. This is the version we think survives you reading the sources yourself.
#Is IELTS marked by a human or by AI?
IELTS Writing and Speaking are marked by trained human examiners, and IELTS Listening and Reading are scored objectively against an answer key. The test owners say so on their own sites. IELTS.org states that every test is assessed by qualified examiners "who undergo rigorous training and continuous monitoring". IDP draws the split explicitly: "The Listening and Reading parts of the IELTS test are marked in a different way to the Speaking and Writing tests." The British Council says that in Writing and Speaking "the examiners use IELTS band descriptors to assess your skills".
The objective half is arithmetic. IELTS.org's scoring guidance says the Reading test contains 40 questions, each correct answer earns one mark, and the total converts to the nine band scale. No judgement is involved, which is why a machine can mark it and always could.
The human half has controls stacked on top. IELTS.org says selected results are marked twice "to ensure that wherever you take your test, your scoring is reliable and consistent", and describes an automatic re-mark where an examiner rated section is out of line with the Listening and Reading scores. That is humans checking humans.
#Does IELTS use AI to mark Writing?
None of the pages IELTS, IDP or the British Council publish on how the test is marked says that AI or automarking is used to score Writing on the live test. The one place an IELTS owner discusses the subject at length is an Insight article on ielts.org published on 13 May 2026, titled "What automarking means for language test validity and integrity". It defines automarkers as algorithms trained with machine learning "to evaluate and mark open-ended spoken and written responses to tasks". It is written generically about high stakes language assessment, it never names IELTS as an adopter, and it frames the exercise as something to understand "before introducing a new scoring system". It is not evidence that IELTS has switched to AI marking.
It does commit to a position. It says "for the time being, a hybrid approach is essential to ensure scoring efficiency and reliability", and closes by saying the priority remains "to use a hybrid model combining automation with a high degree of human oversight and control". It cites the UK regulator Ofqual's guidance that AI cannot be the only means of determining results for high stakes qualifications, and it names the weakness of machine marking plainly: "higher-order skills such as organisation, idea development and the nuances of human communication".
#Is IELTS Speaking marked by AI?
No. IELTS Speaking is a live conversation assessed by a human examiner, and IDP says it in one line: "IELTS Speaking is scored live with a real, human examiner." Elsewhere IDP describes it as evaluated "by certified IELTS examiners in a face-to-face interview using a set of assessment criteria".
Speaking is also the harder module to automate, and the test owner says so. The ielts.org automarking article states that "generally, automarking is more achievable for writing tasks than speaking", because capturing speech, handling accent variation and modelling interaction add layers a written script does not have. Be sceptical of a tool that sounds more confident about Speaking than Writing.
#Why do so many people think IELTS is AI marked now?
Because the tools around IELTS are AI marked, and the two get collapsed into one. Instant AI bands are everywhere in preparation tools now, so it is a short step to assuming the exam works the same way. The ielts.org article draws the line better than we could: "Using automarkers in low-stakes practice contexts has very different implications from a high-stakes university entry test." AI marks your practice. Humans mark your exam. A rehearsal is allowed to be approximate in a way a result is not.
#What does IDP say about using ChatGPT to grade your essay?
IDP tells candidates not to do it. Its guidance on preparing with AI says: "Refrain from using LLM tools to grade your IELTS essays", because the result "will not match up to how human examiners will mark your Writing test, according to official band descriptors".
That is a co-owner of the test advising against products like ours, and we are not going to pretend otherwise. We also think it is largely right about what it describes. A general chatbot has no fixed scoring procedure and no memory of yesterday, so the same essay can come back with a different band on two different afternoons. IDP names the other failure mode too: "AI is well-known for justifying its mistakes and giving overwhelmingly positive feedback." It adds that AI tools "may focus on grammar or pronunciation, but miss nuances of clarity, style or coherence", which is feedback that lists your comma errors and misses a paragraph with no argument in it.
The part we would put on a poster is IDP's workflow rule: "Focus on using AI for correction and feedback, but not creation. Always write your initial answers and draft essays by yourself." Use a scorer to produce the essay and you have practised nothing. There is more on this at can ChatGPT score my IELTS essay.
#How accurate is AI at grading IELTS Writing?
Nobody has published a figure you should treat as settled. The closest thing is Koraishi (2024) in Language Teaching Research Quarterly, which tested ChatGPT 4 on 55 real Writing Task 2 scripts and compared its bands with the official published grades. Agreement was high in aggregate: an intraclass correlation coefficient of 0.814, 95% confidence interval 0.702 to 0.887, and Cohen's weighted kappa of 0.811, interval 0.726 to 0.896. The means came out identical at 6.027, which the author himself calls likely to be misleading and possibly coincidence.
The headline is not the coefficient, it is the conclusion. Koraishi writes that ChatGPT "should not be implemented as an official rater, at least not yet", because individual scripts land well off the mark even when averages line up. Citing 0.814 as proof that AI grades IELTS well inverts his own finding. He also states the limits: 55 essays is a small convenience sample, only Task 2 was tested, and the model was the November 2023 snapshot. There is no like for like human figure to set beside it here. Coefficients from different study designs do not compare, and we are not going to imply otherwise by putting two of them in one sentence.
#What is a practice AI actually for, then?
A practice scorer is a rehearsal marker. It exists to tell you what to fix next, not to predict your result. Two properties matter more than the band it prints.
The first is consistency. Submit the same text to BandNine twice and you get the same band, the same criterion scores and the same feedback, byte for byte, because the first full score is sealed and replayed. A CI check fails our build if that stops being true. Determinism is a consistency property, not an accuracy one. A scorer can be perfectly consistent and consistently wrong. What it removes is the reroll, where you resubmit an unchanged essay until a kinder number appears.
The second is a published figure. An automated run goes out every Saturday at 06:00 UTC through the live production scoring endpoint with the sealed score cache bypassed, and the deviation is published at /quality with the raw JSON alongside it. Read this week's figure there rather than a number in an article. The caveats matter more than the figure. The set is small and its current size is on that page, the labels are written in house by the two of us against the public IELTS band descriptors, and not one script has been marked by a certified IELTS examiner. The figure covers Writing only. We publish nothing equivalent for Speaking, Listening or Reading, and any tool quoting one accuracy number for all four modules deserves a hard look.
#How much should you trust a practice band?
Treat a practice band as a range and a direction, not a verdict. Five rules.
- Read it as plus or minus half a band. If a tool says 6.5, plan for 6.0 or 7.0. Our internal pass mark for a calibration run is a mean absolute error of 0.50, and any script more than a full band out opens a ticket automatically.
- Trust the trend over the digit. Six essays moving from 6.0 to 6.5 to 7.0 tell you something. One essay scored once tells you very little.
- Read the criterion comments, not the headline. The useful output is the sentence about the paragraph with two ideas jammed into it.
- Write first, always. Draft under exam conditions and a clock, then submit.
- Get a human on the final drafts. IDP advises against relying exclusively on AI, and on that we agree. Two or three scripts read by a teacher can catch the organisation and idea development problems that the test owner itself names as the weak spot of machine marking.
For a starting point, take the free diagnostic for a baseline, then work through Task 2 under timed conditions, and read the method behind our own numbers. The short answer stays the same. Your exam is marked by people, your practice is marked by software, and only one of those marks goes on the certificate.