PhonemaBlog

Why Siri and Alexa Don't Understand Your Accent (and What Helps)

Siri, Alexa and voice-to-text misread accented English more often than native speech. Why it happens, which sounds cause most errors, and what actually helps.

Direct answer

Voice assistants are not broken when they mishear you — they were trained mostly on native, often American, English, so accented speech falls outside the sound patterns the model expects. In controlled testing reported by The Washington Post, Google’s and Amazon’s smart speakers were about 30 percent less likely to correctly understand non-American accents than native ones.

That is not a reason to give up on being understood by a machine, or by people. The sounds that confuse a speech-recognition model are usually the same sounds that make a human listener pause or ask you to repeat yourself. Working on them pays off twice.

Why Siri and Alexa mishear accented English

Speech recognition, whether it is Siri, Google Assistant, Alexa, or the dictation keyboard on your phone, works by matching the sound of your speech against patterns learned from huge amounts of training audio. Most of that audio is native English, and a large share of it is American English specifically. The model builds a kind of internal map of which sound is which — where a “t” ends and a “d” begins, how long a vowel needs to be to count as “long,” what a word boundary sounds like in fast speech.

Non-native speech does not sit neatly on that map. A vowel that is slightly shorter than the model expects, a final consonant that is softened or dropped, or a sound produced with a different tongue position can land it in the wrong region of the map entirely. The system is not being unfair on purpose; it is applying statistics learned from a training set that did not include enough speech like yours.

The clearest published evidence for that mechanism does not come from non-native speech at all. A Stanford study published in PNAS tested five major speech-recognition systems and found an average word error rate of about 35 percent for Black speakers of African American English, against about 19 percent for white speakers from the same areas — roughly double the errors, and both groups native speakers of English. The lesson is not that being non-native is the problem. It is that any variety underrepresented in the training audio gets worse accuracy, and non-native English is heavily underrepresented.

The specific sounds that break most often

A few contrasts show up again and again in accent-related transcription errors:

  • The /θ/ sound in think: think, three, and through are frequently transcribed as sink, tree, and threw because /θ/ is rare across the world’s languages and gets replaced with /t/, /s/, or /d/.
  • Word-final consonants: sounds at the end of a word, like the /t/ in want or the /d/ in called, are often softened or dropped in many accents. Dictation software leans heavily on that final sound to pick the right word, so want can become won, and worked can become walk.
  • The short /ɪ/ vowel in ship vs. the long /iː/ vowel in sheep: the length difference between ship and sheep, or live and leave, is small but load-bearing. Speech models rely on it to separate otherwise similar words, and so do human listeners — this is the same contrast behind classic minimal pairs.
  • The /ɹ/ sound in red: most English speech recognition is trained on rhotic (R-pronouncing) varieties. An R produced with a trill, a tap, or a different tongue shape can be misread as a different sound, or dropped from the transcript altogether.

None of these are exotic sounds. They are ordinary, high-frequency parts of English that happen to sit close to a decision boundary the model has to draw somewhere.

Quick fixes for the moment it fails

When a voice assistant or dictation app mishears you right now, a few habits help more than repeating the same sentence louder:

  • Isolate the word it got wrong. Say that one word on its own, then rebuild the sentence around it. A single clear word is easier for both the software and a human to catch than a fast, blended phrase.
  • Give the final consonant its full value. Slightly lengthen the ending sound of the key word — want with a released /t/, called with a clear /d/ — instead of speeding up through it.
  • Add a short pause before the word that matters, not before every word. A rushed sentence with no internal pauses is harder to segment correctly, for a model and for a person.
  • Rephrase instead of repeating. If a specific word keeps failing, a synonym that avoids the confused sound is often faster than fighting the same word five times.
  • Check the transcript, not just your feeling. If a specific pair of words gets swapped repeatedly (think/sink, ship/sheep), that is a concrete, practicable signal — not a vague sense that “my accent is bad.”

The upside: the same practice helps humans too

This is the useful part of an otherwise frustrating problem: none of the sounds above matter only for machines. Final consonants, vowel length, and /θ/ are exactly the kind of details that determine whether a human listener understands you the first time or needs a repeat. If you have ever wondered whether working on pronunciation is really worth it when your accent itself is not a problem, that is a separate and important question — read more on intelligibility vs. accent.

Speech recognition failures are, in a way, an unusually direct and repeatable feedback signal. A person might understand you from context even when a word was unclear; an app has no context to fall back on, so it shows you exactly where clarity broke down.

A newer class of tool takes the opposite approach: instead of trying to understand your accent, it rewrites it on the call in real time. Whether that actually helps depends on whether you want to get through one conversation or be clearer in all of them.

A private way to practice for people, not the cloud

There is a second layer to this problem worth naming: every time Siri, Alexa, or a dictation keyboard mishears you, your voice usually left your device and went to a server to get that answer. If you are already uncomfortable with a system that struggles to understand your accent, sending it your voice repeatedly while you practice is not an appealing trade.

Phonema scores pronunciation entirely on your iPhone, word by word, using the same kind of sound-level detail described above — without uploading your recordings anywhere. You can read more about why on-device processing matters and how the scoring itself works in what GOP (Goodness of Pronunciation) measures.

Bottom line

Siri and Alexa mishear you because your accent is underrepresented in their training audio, not because your English is bad — and the same gap affects native varieties too. The sounds that break the transcript most often are /θ/, word-final consonants, the ship/sheep vowel length contrast, and /ɹ/. Those are the same four things a human listener uses to understand you, which makes them worth practising whether or not a machine is listening.

Related sounds