How AI Scoring Works in TOEFL and IELTS, and Where It Still Fails

TOEFL scores your speaking and writing using a combination of automated scoring and certified human raters, while IELTS still puts a live examiner in the room for speaking. This is the one most important thing that a test taker MUST grasp, as it will alter what each test actually rewards. Machines and people don’t see the same things. If you don’t understand that preparing for automated assessment is a different project, then you are going to end up plateauing when it comes to preparing for a human conversation.

Here is how the scoring actually works in 2026, and where the automation still breaks down.

How AI Scoring Works in TOEFL and IELTS, and Where It Still Fails

The scoring architecture

Objective sections are easily programmable. Reading and listening have right answers and, for decades, have been machine-marked, and nothing interesting is occurring there.

It’s the difficult part that involves productive skills.

The ETS scores TOEFL speaking and writing responses with a blend of certified human and AI scoring. Neither works alone. The automated system gives a score, the human gives a score, and the design is based on the assumption that each will get it that the other doesn’t.

IELTS can handle speaking variations. It’s still a live interview with an interviewer that can interrupt, follow up, and redirect depending on the candidate’s answers.

This is not a quality ranking. It’s a real design compromise. Automated scoring is both consistent and quick, not a 4 pm Friday fatigue artist. A human examiner can react to something unexpected. Both are purchasing something that the other one does not have.

What changed in 2026

Worth flagging because a lot of preparation material has not caught up: TOEFL moved to a 1 to 6 scale on 21 January 2026. Candidates now receive four section scores and an overall score, which is the average of the four, rounded to the nearest half band. A comparable 0-to-120 score continues during a two-year transition period.

ETS publishes the current section structure and scoring directly. Any practice material describing a 0 to 120 scale as the primary output predates the change.

The test is also now two hours long, and scores are available in approximately 3 days.

Working out which criterion is actually holding a candidate back usually needs someone trained in the descriptors, which is why targeted IELTS and TOEFL preparation still outperforms self-study for anyone stuck half a band below target.

What automated scoring is good at

Consistency. No matter when the response is written, or if others are being evaluated on the same day, the same score will be given to the response. Human raters drift. Machines do not.

Measurable surface features. Lexical range, grammatical accuracy, sentence complexity, pause patterns, and speech rate. The systems work well with these, and they are truly measurable.

Detecting the obviously off-task. A response that does not address the prompt gets caught reliably.

Scale. Millions of responses checked off in days, the reason the technology has been invented.

Where it still fails

This is the part worth understanding if you are preparing.

Content quality is approximated, not assessed. It can mark a well-structured, linguistically rich, and appropriately length response or not. Whether the argument is any good is another question, and current systems only deal indirectly with it at best. A phrase(s) delivered fluently and organized, but trivial, scores better than it deserves.

Accented speech remains a known weak point. The automated speech scoring has not been as consistent across accents as it has been across proficiency levels. This is something providers work on continually, and part of the reason for using human raters is because of it; but it’s fair to say that if a candidate has a strong regional accent, they should want a format with the human element.

Unusual but correct responses get penalized.  The distributions of previous responses have been used to train automated scoring. If the answer is outside of the training distribution, but is structurally atypical, this is not an excellent answer. Conventional competence is a slightly depressing thing, but it is the safe bet with automated scoring.

Interactive skill is not measured at all. A recorded monologue is not the best way to determine if you can be interrupted, clarify, or recover if a conversation goes off course. Those are some of the most practical real-world skills for speaking, and TOEFL’s format can’t get them.

What this means for preparation

The strategies genuinely diverge.

For TOEFL speaking, practice producing complete, well-structured responses into a microphone within a fixed window with no feedback. Structure matters more than personality. Signposting matters more than charm. Get comfortable with the absence of a listener, which most candidates find harder than the language itself.

For IELTS speaking, practice conversation with a person who will interrupt and follow up. Recording yourself alone will not prepare you for an examiner changing direction mid-answer.

For writing on either, structure and development are overemphasized. The clear organization is rewarded by both automated systems and human raters, making it a part of the system that can be learned the easiest.

A note on AI preparation tools. Language models can actually be used to create an infinite number of practice questions and describe grammar. They are not suitable for predictions of band scores, es as descriptors are holistic judgments and the models are intended to give confident numbers that change when asked a second time. Do not use them to calibrate the meter.

The honest summary

Automated scoring in language testing isn’t either the language testing monster that its detractors say it is, or the solved problem suggested by its proponents. Excellent with consistency and measurable surface characteristics; weak on the quality of content, unusual responses, and anything interactive.

The prudent course: Don’t avoid tests involving it. Recognize that TOEFL will give points for a well-structured, conventional, clearly signposted response, while IELTS will give points for a real conversation.

Choose the format that is the most appropriate for your performance, and practice that format, not English in general.

Popular on OTW Right Now!

Add a Comment

Your email address will not be published. Required fields are marked *