
AI agents do worse outside English: how to check yours
An August 2026 study found every AI agent it tested did worse outside English, mostly at actually doing the job. What that means for Mandarin and Cantonese callers.
Say a dental clinic uses an AI agent on its phone line: software that answers for you and can act for you, like moving an appointment. A mother calls and asks, in Mandarin, to move her son's check-up to Thursday afternoon. The agent replies in fluent Mandarin and reads out Thursday's free times. Then it ends the call without moving the booking. Or it moves it twice. The call sounded fine. The appointment book is wrong.
That call is made up. The pattern is not. A study published in August 2026 measured this kind of failure in seven AI agents working outside English.
When we read them in September 2026, the voice company ElevenLabs' model documentation said one of its speech models "supports 70+ languages", and its agent product page offered to "Support customers in 70+ languages". That model's list named Mandarin Chinese, not Cantonese. A number like that describes the voice, not whether the AI behind it books the right patient at the right time.
What the August 2026 study found
The study, by a team that included Meta researchers, gave AI agents multi-step jobs in everyday apps such as email, calendar, contacts and shopping, in English and ten other languages. Every agent did worse outside English. Of the five languages other than English most used at home in Australia's 2021 Census (Mandarin, Arabic, Vietnamese, Cantonese and Punjabi), only Mandarin was tested.
The full paper shows where things went wrong:
- Doing, not understanding. The agents handled numbers about as well as in English. They got worse at actually making changes: too few or too many. They also hesitated longer before acting. Their final replies got worse too.
- Mostly the AI's own fault. On the authors' numbers, 55.4% of the drop was the AI getting it wrong. The rest came from flaws in the study's own translation and automatic marking.
- Trying again did not help. For most agents, the same tasks in the same languages failed every time.
The drop was biggest in Mandarin, Japanese and Hindi, the three tested languages not written in the English alphabet. In an appendix on one agent, the most common mistake of this kind was searching for a name in one form, such as Chinese characters, when the records held it in another, such as English letters. Often the agent noticed both forms and still searched for only one. Picture records that say Chan Tai Man, and a caller who says 陳大文.
A bigger AI did not close the gap
The paper tried four sizes of one family of AI models. Scores rose with size, but English stayed on top every time. At the biggest size, English scored 31.6 and Mandarin 16.2. Telling that model to think harder did not lift its score either.
That is one family in one study, not proof that no bigger AI is better in Mandarin. But size alone did not remove the difference, so an upgrade is no replacement for testing.
Cantonese is harder again
Cantonese was not in the study. Other research points the same way. A paper posted on 7 September 2026 notes that Cantonese is widely spoken but has little written material for AI to learn from. When the team trained a model to reason in Cantonese, one stage, using a little Cantonese reasoning data, made it score worse. A later stage won that back.
Our reading, not the paper's: a little Cantonese data, used the obvious way, made the AI worse. So "tuned for Cantonese" is a claim to test, not a reason to skip testing.
On 30 October 2025, the Chinese University of Hong Kong announced a platform for testing Cantonese AI, saying "even the latest models still struggle to fully capture the nuances of Cantonese". It tests everyday written Cantonese and the way speakers mix in English and Mandarin, which is what a phone agent hears. 我想 book 個 appointment 洗牙 (I want to book a teeth-cleaning appointment) is one sentence in two languages.
How to check your own agent
Published AI scores will not test this for you: the study's authors note that other-language testing of leading AI agents is mostly one-off quiz questions, when it is done at all. A first check fits in an afternoon. Write ten requests the way your Mandarin or Cantonese callers actually say them. Run each one three times. Compare what the agent entered with what your receptionist would have. A fluent answer with no booking is a fail. So is the same booking made twice. Include:
- Names in two forms. Records say Chan Tai Man; the caller says 陳大文. The agent should find the existing client, not create a new one.
- Spoken amounts. Chinese counts big numbers in 萬, units of ten thousand. A mortgage caller asking about 八十萬 means 800,000.
- Clock times. In Cantonese, 字 is also five minutes, so 六點兩個字 is 6:10. A booking at 6:00 sounds fine on the call and is still wrong.
- Mixed languages. Cantonese with English words in it, like the sentence above.
- Changes that must happen. Reschedule, cancel, or log an urgent after-hours plumbing job. Check it was done, and done once.
What to ask your provider
- In which languages has the agent been tested on doing the job, not just speaking? Can I see the results?
- Does it handle Cantonese, or only Mandarin?
- When it is unsure what it heard, does it read the name, amount and time back to the caller, in their language, before saving anything? If that fails, does the call go to a person?
- Can I see what the agent entered in my system after each call, not just a transcript?
- Will you tell me before you change the AI, its instructions or the voice, so I can run my ten test calls again?