How Accurate Is AI Transcription? What the Numbers Mean
Short answer: on clean audio — one person, close to a decent microphone, quiet room — the better transcription tools now get roughly 95 to 98 percent of words right. On the recording you actually have, which is three people on a video call and someone dialling in from a car, expect somewhere between 80 and 92 percent. The first number is the one on the vendor home page. The second one decides whether the transcript saves you any time.
What 95 percent accurate costs you
Transcription accuracy is the flip side of word error rate, usually written WER. It counts three kinds of mistake against the number of words actually spoken: words swapped for the wrong word, words dropped, and words the tool added that nobody said. Ninety-five percent accurate means a WER of about five percent.
An hour of ordinary conversation runs to something like 8,000 spoken words. Five percent of that is 400 wrong words, or about one every nine seconds. Spread evenly, that would be survivable. It is not spread evenly. Speech recognition is most confident about common words and least confident about rare ones, so the errors collect exactly where the information is: names, figures, product names, acronyms, and whatever gets said fast at the end of a sentence.
That is what the percentage really means. A 95-percent transcript is not 95 percent as useful as a perfect one. It is a transcript where most of the sentences are fine and most of the facts need checking.
What actually sets your number
The recording matters more than the tool. Roughly what to expect:
| What you recorded | Typical words correct |
|---|---|
| One speaker, close mic, quiet room | 95–98% |
| Recording of a video call | 85–92% |
| Phone audio | 80–88% |
| Noisy room or café | 70–85% |
| Strong or unfamiliar accent | 75–90% |
Those ranges come from AssemblyAI, summarising its own July 2026 benchmarking. AssemblyAI sells speech-to-text, so read them as a supplier describing a good day rather than an independent test. The shape is the useful part: fifteen to twenty points separate the best case from the ordinary one, and switching tools does not close that gap.
The same audio, different speakers
Accuracy is not evenly distributed across people either. A 2020 study in PNAS put matched interview recordings through the speech recognition systems of Amazon, Apple, Google, IBM and Microsoft and found an average word error rate of 0.35 for Black speakers against 0.19 for white speakers — close to double the errors on comparable audio. Those systems have been retrained many times since and nobody should assume the same figures hold today. The mechanism has not changed, though. A model is accurate on the speech it heard a lot of in training and less accurate on everything else, so if your recordings involve accents, dialects or second-language speakers, the headline number is not your number.
I asked ChatGPT to clean up a bad transcript
Cleaning up the raw output is the next step most people reach for, so I tested it. I wrote a ten-line transcript of an invented bakery staff meeting and seeded it with the mistakes transcription actually makes: a homophone (flower where the speaker meant flour), an affect-for-effect swap, a garbled figure (1,40 croissants), a mangled name (Ronnie a), an [inaudible] gap, and two lines that contradict each other about which day a repair is happening. Then I asked ChatGPT to clean it up without changing any facts and to list every change it made.
Most of it went better than expected. It corrected flower to flour from the bakery context and fixed the affect-for-effect swap. More importantly, it refused to guess at what it could not know: it left 1,40 croissants exactly as written and said why, left two 3 weeks alone, kept the [inaudible] marker rather than inventing words to fill the gap, and called out the contradiction about the repair day instead of quietly choosing one.
One change it did not flag as a guess: Ronnie a came back as Ronnie A., listed under formatting. That is not formatting. Two words the model could not place became a person with a middle initial — a name that now looks deliberate and gets copied into the minutes by whoever reads it next. The single thing it invented was the single thing it had no way of knowing.
Cleanup is worth doing, and asking for the change list is worth doing. Read that list against the transcript rather than taking its word for what it did. The same habit applies to anything an AI hands back; there is a quick method in how to fact-check an AI answer in under two minutes.
Measure your own in about ten minutes
Published benchmarks are run on tidy datasets. Your figure is the only one that matters, and you can get it over a coffee.
- Take two minutes of the kind of audio you actually deal with — the weekly call, not your best microphone.
- Run it through the tool you are considering.
- Play it back and type what was really said. Two minutes is about 270 words.
- Compare the two and count three things: words that came out wrong, words that went missing, and words that appeared from nowhere.
- Add those three numbers and divide by the number of words actually spoken. That is your word error rate.
Under five percent, you can skim and fix as you read. Between five and ten, budget a real editing pass before the transcript goes anywhere. Above ten percent, you will finish sooner working from the audio and using the transcript only to find your place.
Where the extra accuracy comes from
Almost all of it comes from the recording, not the software. Get the microphone closer to whoever is talking. Record locally instead of relying on the meeting platform compressed stream. Ask people not to talk over each other, and have everyone say their name once at the start — it gives you a clean reference for every mangled version that follows. If the tool accepts a custom vocabulary list, feed it your product names and jargon before the first run rather than fixing them fifty times afterwards.
After that, the tool choice is worth a few points. Whisper and the newer hosted models handle accents better than the older platform captions do, and Descript is built around editing the errors rather than just reporting them. Which of them fits your work is the subject of the comparison below.
Related reading
- Otter vs Rev vs Descript vs Whisper — which tool to reach for once you know what your audio is like.
- AI Meeting Note-Takers: Otter vs Fireflies vs Fathom — for recurring meetings where the summary matters more than the transcript.
- How to Turn Your Rough Notes Into a Polished Document With AI — what to do with the transcript once it is clean.
Written by Mitch, a software analyst who tests software for a living. Every guide here comes from actually using the tool on a real task — including the parts where it falls over. Tested on ChatGPT Plus, Claude Max and Gemini (free). More about this site · Corrections: [email protected].
Get the next guide by email
A new tested how-to when there is one — usually a couple a month. No hype, unsubscribe any time.
Tested in real accounts. No affiliate links.