Why Is Voice to Text So Bad? Causes and Fixes
Why is voice to text so bad? (Or talk to text, or voice typing: same engines, same question.) Mostly because speech recognition is a chain (microphone, room, voice, vocabulary, model) and your transcript is only as good as the weakest link. You said “their” and got “there”. You said your manager's surname and got a vegetable. You stopped to think, and a sentence appeared that you never said.
Each has a different cause and fix. Below: symptoms, causes, fixes, then a straight answer to whether voice to text counts as AI. Shopping for a tool instead? Start with our dictation software roundup.
Voice to text fails for a handful of predictable reasons
Voice to text gets words wrong when the audio is unclear, when the words are ones the model rarely met in training, or when it lacks the context to choose between words that sound alike. Most complaints trace back to one of those three, and the symptom usually tells you which.
| What you see | Likely cause | Try first |
|---|---|---|
| Random words wrong, worse in some rooms | Mic too far away, or an echoey or noisy room | Bring the mic within a hand's width of your mouth; check which mic is selected |
| Fine at your desk, hopeless outdoors | Wind and traffic hitting the phone's mic | Cup the phone or use earbuds with an in-line mic |
| Names, brands and acronyms always wrong | Words the model rarely saw | A custom dictionary, or correcting by voice |
| “There” for “their”, “two” for “to” | Sound-alikes settled by context | Speak in whole phrases; proofread for these on purpose |
| First word or two missing | Talking before it starts listening | Wait for the listening cue |
| Phrases you never said, after a long pause | The model filling silence | Pause the mic while you think |
| Gibberish, or the wrong language | Wrong dictation language or region | Set the language and the closest regional variant |
| It suddenly got worse | Something in the chain changed | Retrace: new earbuds, update, new keyboard, offline pack |
The microphone matters more than the model
Poor audio is the first thing to rule out when voice to text is bad: a microphone far from your mouth records the room nearly as loudly as your voice, and no model can recover words that never arrived clearly. Laptop microphones sit beside fans and keyboards. A phone held at arm's length on a busy street is mostly recording the street.
Headsets aren't automatically better. Many Bluetooth earbuds drop into a lower-quality call mode while their mic is live, so a cheap wired pair can beat expensive wireless ones. And check what the computer is listening to: Mac Dictation, Windows voice typing and Word's Dictate each let you pick the microphone, and a webcam mic across the room can quietly become the default.
The vendors agree. Microsoft's troubleshooting for Word's Dictate suggests a headset or external mic and less background noise; Google's voice typing help says move somewhere quiet and plug in an external microphone.
Hear what the software hears: record twenty seconds in your phone's voice memo app, from exactly where you normally dictate, then play it back on headphones. If you can hear the dishwasher, the extractor fan or your own keyboard, so can the recognizer, and moving the mic will help more than changing apps.
How you talk changes what it hears
Speaking style affects accuracy because recognizers learn largely from natural, fluent speech, so trailing off, swallowing the ends of sentences or dictating one slow word at a time gives the model less to work with. Microsoft's advice to “speak more deliberately” holds if deliberate means complete, not slow: think the sentence through, then say it in one breath.
Long pauses carry a stranger risk. In a 2024 study titled “Careless Whisper”, Allison Koenecke and colleagues found that roughly 1% of transcriptions from OpenAI's Whisper contained whole phrases or sentences that weren't in the audio, and that these hallucinations happened disproportionately for speakers with longer stretches of silence. Most dictation tools aren't built on that exact model, but the lesson carries.
Think with the mic off: when you need more than a few seconds to find the next sentence, stop dictation, think, then start again with the whole sentence ready. You get fewer half-finished phrases, and the engine never has a long silence to fill.
Names, jargon and homophones need context it doesn't have
Voice to text struggles with names and jargon because a recognizer leans towards words it met often in training, and it chooses between sound-alikes like “their” and “there” using context it frequently lacks. A phone keyboard hears one sentence in isolation. It doesn't know you're replying to Aoife about the Q3 renewal, so it reaches for the commonest words that fit the sounds.
For rare words, hand the tool your vocabulary. Dedicated AI dictation apps have custom dictionaries; the Laxis Voice Keyboard calls its version a Personal Dictionary, for names, acronyms and technical terms. Built-in dictation mostly lacks one, so correct in place instead: click a blue-underlined word on a Mac for alternatives, say “change” and spell the fix on an iPhone, or tap the misheard word on a Pixel. Google says Pixel's advanced voice typing improves from the words you correct.
Homophones need a different habit. Longer phrases give the language side of the model more to go on, so “send it to their office” comes out better than “their” spoken alone. Then proofread specifically for sound-alikes, numbers and negatives, the three places a single wrong word flips the meaning.
Accents and dialects are unevenly served
Voice to text is measurably less accurate for some accents and dialects, because the training data behind most engines over-represents some ways of speaking. The best-known evidence is a 2020 Stanford study in PNAS led by Allison Koenecke: systems from Amazon, Apple, Google, IBM and Microsoft, tested in 2019, misrecognised an average of 35% of words from Black speakers against 19% from white speakers. Engines have improved since, but nobody has shown the gap is closed.
Choosing the regional variant closest to how you speak helps, and it costs nothing. We go further in our guide to dictating with an accent or in two languages.
Phone keyboards run small models; cloud services can run big ones
Phone dictation often runs on the device itself, where the speech model must fit within a phone's memory and battery budget, while cloud services can run far larger models. Apple says iPhone dictation is processed on the device in many languages with no internet needed. Google says Pixel's advanced voice typing keeps what you say on the phone unless you use features like Fix it. Windows voice typing, by contrast, uses online recognition powered by Microsoft's Azure Speech services.
On-device isn't worse by definition. It's private, it works on a plane, and phone makers are adding local language models to tidy the result; in iOS 27, Apple uses one to improve spelling, punctuation and capitals in English on iPhone 17 Pro, iPhone Air and later models. But it's a trade-off, and one reason the same sentence can come out differently on phone and laptop; our explainer on on-device versus cloud transcription covers the privacy side.
If voice to text feels like it's getting worse, the technology probably hasn't regressed; something in your chain has changed, such as new earbuds, a new default keyboard, an OS update or an offline language pack. Retest close to the mic in a quiet room to see whether the problem follows the audio or the software.
Is voice to text AI?
Voice to text is AI in the machine-learning sense: modern speech recognizers are neural networks trained on large amounts of recorded speech paired with transcripts, not sets of hand-written rules. So yes, it counts. OpenAI's Whisper, described in a 2022 paper by Alec Radford and colleagues, was trained on 680,000 hours of audio collected from the web. Every mainstream dictation engine today is built this way.
That explains most of the behaviour above. The model predicts the likeliest words for what it heard, based on what it learned, so it's strong on common speech and weaker on rare names and under-represented voices. It can be confidently wrong, and occasionally produce words nobody said. Our speech to text explainer goes deeper on how recognizers work.
What it doesn't mean is that dictation writes for you. Basic voice to text transcribes; it isn't generative AI in the chatbot sense. The line is blurring, because many tools add a second step that edits your words with a language model: Windows Fluid dictation, Pixel's Fix it, Apple's iOS 27 model and Laxis's Verbal Cleanup all do versions of this. If a school or employer policy restricts “AI”, ask whether it means transcription or rewriting, then check which extras are switched on.
When built-in dictation is good enough
Built-in dictation is good enough for short messages, searches and quick notes in a quiet room, and it's free. Fix the microphone, language setting and pauses first. If what's left is mostly names, filler words and rambling sentences, an AI dictation app earns its keep.
Laxis Voice Keyboard works across your desktop and phone apps, removes filler and repetition, punctuates, and uses your Personal Dictionary for names. The limits: it's cloud-only, so it needs a connection and your audio leaves the device, and on the free plan, dictation and meetings draw on one shared 300-minute monthly pool. The 14-day Premium trial needs no credit card, so you can test it against your own worst errors.
Frequently asked questions
Why is voice to text so bad on iPhone?
Voice to text on iPhone usually goes wrong because of microphone distance, background noise, or names and jargon it has never learned, rather than a flaw unique to the iPhone. Hold the phone closer, dictate in complete phrases, and spell hard names letter by letter. On iPhone 12 and later in U.S. English, you can also fix a misheard word by voice with 'change'.
Why is voice to text getting worse?
Voice to text usually seems to get worse because something in your setup changed, not because the technology regressed. Common culprits are new Bluetooth earbuds, a different default microphone, an operating system update, a new keyboard app or an offline language pack. Retest with the microphone close in a quiet room to see whether the problem follows the audio or the software.
Is voice to text AI?
Yes. Modern voice to text uses machine learning: a neural network trained on large amounts of recorded speech and matching transcripts predicts the most likely words for the sound it hears. OpenAI's Whisper, for example, was trained on 680,000 hours of audio. Plain dictation transcribes rather than writes, but many tools now add an AI step that tidies or rewrites the text.
Is speech to text generative AI?
Basic speech to text is not generative AI in the chatbot sense, because it transcribes what you said instead of creating new content. The line is blurring: recognizers produce text word by word and can occasionally insert phrases nobody spoke, and many dictation tools add a language model that rewrites your words. Check which features are on if a policy restricts generative AI.
How can I make voice to text more accurate?
Move the microphone closer to your mouth, cut background noise, speak in complete phrases at a normal pace, and choose the right language and regional variant. Then teach the tool your vocabulary, using a custom dictionary where it has one. Those steps fix most errors; if names and filler words are still the problem, try a dedicated AI dictation app.
Why does voice to text get names wrong?
Voice to text gets names wrong because recognizers favour words they saw often in training, and most surnames, brands and acronyms are rare. Without context, the engine substitutes the nearest common word. Adding the name to a custom dictionary fixes it on tools that have one; built-in dictation mostly doesn't, but spelling the name letter by letter works on iPhone and newer Pixels.