You pasted one essay into two AI scorers. One came back 6.5, the other 7.5. Nothing changed between the two clicks, so at least one number is wrong, and neither screen tells you which. The causes are mechanical, and you can diagnose most of them yourself.
#Why do two AI tools give the same essay different band scores?
Two AI tools give the same essay different band scores because the band is not read off your text, it is generated by a model, and four things move it: the randomness setting the tool runs at, which version of the model is answering today, the instructions and rubric wrapped around that model, and whether the tool was shown the task question at all.
#1. Sampling temperature
A language model does not pick the single most likely next token every time. Temperature controls how much variation is allowed when it chooses. Low, and the output is nearly fixed. Higher up, the same input produces a different chain of reasoning, and that can land on a different band. A tool left at a chatty default wobbles on identical input.
#2. The model underneath is not pinned
Model providers ship new versions and retire old ones. A tool that calls whatever the current model is, rather than a pinned version, changes its marking behaviour the day the provider updates. Your March essay and your September essay were marked by two different engines, and nobody told you.
#3. The prompt and rubric wrapped around the model
This is the least visible of the four. The public band descriptors are the same document for everyone. What differs is what each tool does with them: whether it marks the four criteria separately or gives one holistic number, how it rounds, whether it is told to be strict or encouraging, whether it has worked examples at each band to anchor against. Two competent builds of the same rubric can land on different bands and both be defensible.
#4. Whether the tool was given the task question
An essay marked without the question cannot have Task Response marked properly, because Task Response is a judgement about how well you answered that prompt. If it received only your essay, the model reverse engineers the question from your answer. An essay that drifts off the real task then looks fine, because the inferred question is the one it answers.
The check is quick. Look for a field where the task prompt goes. If there is not one, anything the tool says about Task Response is guesswork, and Task 1 is worse: without the chart or letter prompt, no accuracy judgement is possible. Our Task 2 guide sets out what each criterion wants.
#What is a band score actually measuring?
A band is a descriptor judgement reported in half-band steps, not a percentage, so the smallest disagreement that can exist between two markers is 0.5. A 6.5 against a 7 is one step, the narrowest gap available, and can come from two markers who read your essay the same way and split on one criterion. A 6.5 against a 7.5 is two steps, which means they are reading it differently. Our band score explainer covers how the four criteria combine.
#Is one tool giving you two different scores worse than two tools disagreeing?
Yes, and it is much worse. Two tools disagreeing tells you they were built differently, which is expected. One tool giving two answers for the same text tells you it has no settled opinion, which means the first number was partly a coin toss. Between products, disagreement is a calibration problem. Inside one product, it is a reliability problem.
The damage lands on progress tracking. If a scorer swings half a band on identical input, a real half band of improvement is invisible inside its own noise, and you cannot tell a better essay from a luckier roll.
#How do you test a scorer for consistency in two minutes?
Submit exactly the same text twice and compare everything, not just the overall band. Score it, wait, paste the identical text and score again. Check the four criterion scores and the wording of the feedback too. A tool that shifts Coherence and Cohesion from 6.5 to 7 while the overall band holds steady is still unstable, it has hidden the movement in the average.
On BandNine that test has a fixed answer. The first full score for a piece of text is sealed, and every later submission of that same text against that same question returns it byte for byte: same overall band, same criterion scores, same feedback sentences. An automated check in our build enforces it, so a change that broke it would fail CI. Try to break it.
Be clear about what that is worth. Determinism is a consistency guarantee, not an accuracy one. A scorer that is consistently half a band generous is consistently wrong. Consistency buys one thing: when the number moves, you know your writing moved it.
Accuracy is measured separately, and we go into it further in how accurate AI IELTS scoring is. The run is automated every Saturday at 06:00 UTC, through the live production scoring endpoint with the sealed-score cache bypassed, over 24 scripts labelled in house by the two of us against the public band descriptors. None has been marked by a certified IELTS examiner, and the endpoint says so in its own output. Our internal pass mark is a mean absolute error of 0.50, and anything worse, or any script out by more than a full band, opens a ticket. This week's figure and the raw JSON sit on the quality page. That number covers Writing. We do not publish one for Speaking, Listening or Reading.
#Does the research say AI band scores can be trusted?
The most useful published study finds strong agreement in aggregate and still concludes that AI should not be an official rater. Koraishi (2024), in Language Teaching Research Quarterly, put 55 real Writing Task 2 scripts through ChatGPT 4 as it existed in November 2023 and compared its bands with the published official scores. Agreement was substantial: an intraclass correlation coefficient of 0.814, 95% confidence interval 0.702 to 0.887, and a Cohen's weighted kappa of 0.811. The two sets of means matched exactly at 6.027, which the author calls likely misleading and possibly coincidence. His conclusion is that ChatGPT should not be an official rater yet, because individual scripts still come out well off despite the aggregate fit. The sample is small, covers Task 2 only, and is tied to one model snapshot.
For scale, IELTS publishes reliability statistics for its own marking, and has Writing and Speaking marked by trained examiners with selected scripts marked twice. The study designs are not equivalent, so the two sets of numbers cannot be read side by side. If a chatbot handed you a band, this page covers that case directly.
#What do the test owners say about AI marking?
IDP, one of the IELTS co-owners, tells candidates not to use large language models to grade their essays. Its guidance on AI in preparation says "Refrain from using LLM tools to grade your IELTS essays", because the result will not match how human examiners mark against the official band descriptors. The same page recommends AI for correction and feedback but never for creation, and warns that AI feedback skews positive, catching grammar while missing clarity and coherence.
IELTS itself, in a May 2026 insight article on automarking, reports that many experts now advocate a hybrid, where any response the machine is not confident about is escalated to a human examiner. It cites the UK regulator Ofqual's guidance that AI cannot be the only means of determining results for high-stakes qualifications, and names higher-order skills such as organisation, idea development and nuance as the weakness of automarkers. It is forward looking. It does not say IELTS marks the live test by machine, and none of their pages explaining how marking works says so: Writing and Speaking are marked by trained human examiners, with selected results marked twice. It draws one distinction that matters here: using automarkers in low-stakes practice, it says, has very different implications from a high-stakes entry test.
#What should you do with two conflicting band scores?
- Treat every single band as an estimate. One score from one tool on one essay is a reading, not a result. Direction across several essays from one stable scorer is worth more than any single number.
- Compare the feedback, not the numbers. If both tools flag the thin second body paragraph, the missing overview and the same article errors, they agree about your essay and disagree only about the label. Fix what they agree on.
- Prefer the tool that shows its working. Four criterion scores, sentences quoted from your own text, a reason tied to a descriptor. A tool that returns a bare number is not marking, it is guessing in public.
- Give it the question. If there is no field for the task prompt, discount the Task Response line entirely.
- Run the same-text test before you trust a trend. Two minutes, and you know whether it can measure change at all.
- Get a human in the loop for the essays that matter. IDP's own advice is to mix AI feedback with human feedback rather than rely on one.
The honest answer to "which tool is right" is usually neither, precisely. Both estimate a descriptor judgement a human marker makes on test day, with different settings behind the glass. Use the number for direction and the feedback for action. Our method is on the quality page, and you can run the same-text test on us with the free diagnostic. Do both attempts in one sitting, against the question it gives you, because a new question counts as a new submission.