How Accurate Is AI Transcription? What the Numbers Mean

How Accurate Is AI Transcription? — The AI Cheat Sheet

Short answer: on clean audio — one person, close to a decent microphone, quiet room — the better transcription tools now get roughly 95 to 98 percent of words right. On the recording you actually have, which is three people on a video call and someone dialling in from a car, expect somewhere between 80 and 92 percent. The first number is the one on the vendor home page. The second one decides whether the transcript saves you any time.

What 95 percent accurate costs you

Transcription accuracy is the flip side of word error rate, usually written WER. It counts three kinds of mistake against the number of words actually spoken: words swapped for the wrong word, words dropped, and words the tool added that nobody said. Ninety-five percent accurate means a WER of about five percent.

An hour of ordinary conversation runs to something like 8,000 spoken words. Five percent of that is 400 wrong words, or about one every nine seconds. Spread evenly, that would be survivable. It is not spread evenly. Speech recognition is most confident about common words and least confident about rare ones, so the errors collect exactly where the information is: names, figures, product names, acronyms, and whatever gets said fast at the end of a sentence.

That is what the percentage really means. A 95-percent transcript is not 95 percent as useful as a perfect one. It is a transcript where most of the sentences are fine and most of the facts need checking.

What actually sets your number

The recording matters more than the tool. Roughly what to expect:

What you recordedTypical words correct
One speaker, close mic, quiet room95–98%
Recording of a video call85–92%
Phone audio80–88%
Noisy room or café70–85%
Strong or unfamiliar accent75–90%

Those ranges come from AssemblyAI, summarising its own July 2026 benchmarking. AssemblyAI sells speech-to-text, so read them as a supplier describing a good day rather than an independent test. The shape is the useful part: fifteen to twenty points separate the best case from the ordinary one, and switching tools does not close that gap.

The same audio, different speakers

Accuracy is not evenly distributed across people either. A 2020 study in PNAS put matched interview recordings through the speech recognition systems of Amazon, Apple, Google, IBM and Microsoft and found an average word error rate of 0.35 for Black speakers against 0.19 for white speakers — close to double the errors on comparable audio. Those systems have been retrained many times since and nobody should assume the same figures hold today. The mechanism has not changed, though. A model is accurate on the speech it heard a lot of in training and less accurate on everything else, so if your recordings involve accents, dialects or second-language speakers, the headline number is not your number.

I asked ChatGPT to clean up a bad transcript

Cleaning up the raw output is the next step most people reach for, so I tested it. I wrote a ten-line transcript of an invented bakery staff meeting and seeded it with the mistakes transcription actually makes: a homophone (flower where the speaker meant flour), an affect-for-effect swap, a garbled figure (1,40 croissants), a mangled name (Ronnie a), an [inaudible] gap, and two lines that contradict each other about which day a repair is happening. Then I asked ChatGPT to clean it up without changing any facts and to list every change it made.

ChatGPT returning a cleaned-up meeting transcript, having corrected flour but left the garbled number and inaudible marker in place
The cleaned transcript. ChatGPT, September 2026, default model selected.

Most of it went better than expected. It corrected flower to flour from the bakery context and fixed the affect-for-effect swap. More importantly, it refused to guess at what it could not know: it left 1,40 croissants exactly as written and said why, left two 3 weeks alone, kept the [inaudible] marker rather than inventing words to fill the gap, and called out the contradiction about the repair day instead of quietly choosing one.

One change it did not flag as a guess: Ronnie a came back as Ronnie A., listed under formatting. That is not formatting. Two words the model could not place became a person with a middle initial — a name that now looks deliberate and gets copied into the minutes by whoever reads it next. The single thing it invented was the single thing it had no way of knowing.

Cleanup is worth doing, and asking for the change list is worth doing. Read that list against the transcript rather than taking its word for what it did. The same habit applies to anything an AI hands back; there is a quick method in how to fact-check an AI answer in under two minutes.

Measure your own in about ten minutes

Published benchmarks are run on tidy datasets. Your figure is the only one that matters, and you can get it over a coffee.

  1. Take two minutes of the kind of audio you actually deal with — the weekly call, not your best microphone.
  2. Run it through the tool you are considering.
  3. Play it back and type what was really said. Two minutes is about 270 words.
  4. Compare the two and count three things: words that came out wrong, words that went missing, and words that appeared from nowhere.
  5. Add those three numbers and divide by the number of words actually spoken. That is your word error rate.

Under five percent, you can skim and fix as you read. Between five and ten, budget a real editing pass before the transcript goes anywhere. Above ten percent, you will finish sooner working from the audio and using the transcript only to find your place.

Where the extra accuracy comes from

Almost all of it comes from the recording, not the software. Get the microphone closer to whoever is talking. Record locally instead of relying on the meeting platform compressed stream. Ask people not to talk over each other, and have everyone say their name once at the start — it gives you a clean reference for every mangled version that follows. If the tool accepts a custom vocabulary list, feed it your product names and jargon before the first run rather than fixing them fifty times afterwards.

After that, the tool choice is worth a few points. Whisper and the newer hosted models handle accents better than the older platform captions do, and Descript is built around editing the errors rather than just reporting them. Which of them fits your work is the subject of the comparison below.

Related reading

Written by Mitch, a software analyst who tests software for a living. Every guide here comes from actually using the tool on a real task — including the parts where it falls over. Tested on ChatGPT Plus, Claude Max and Gemini (free). More about this site · Corrections: [email protected].

Get the next guide by email

A new tested how-to when there is one — usually a couple a month. No hype, unsubscribe any time.

© 2026 The AI Cheat Sheet  ·  About  ·  Contact  ·  Privacy Policy  ·  RSS
Tested in real accounts. No affiliate links.

Popular posts from this blog

Is It Safe to Paste Work Documents Into ChatGPT?

How to Use AI to Write and Tailor Your Resume and Cover Letter

How to Spot AI Scams and Deepfakes Before They Cost You