Start Free Trial
← All posts
Reading the fine print

Word error rate, explained

Every accuracy percentage you have ever seen on a transcription product is a WER measurement in disguise. Here is how the math works, and how to read the claims like someone who knows.

3 error types 8 min read Published August 2026

"highly accurate." "Industry-leading accuracy." "Near-perfect transcription." Every one of these claims comes from the same underlying measurement — word error rate — and once you know how it is computed, you can tell which claims mean something and which are marketing weather. Accuracy is not a property of an engine; it is a property of an engine plus the audio you feed it — which is why the only number that should decide your purchase is the one you measure on your own recordings. The math takes two minutes to learn.

The formula, with a worked example

WER compares what the engine wrote against what was actually said, counting three kinds of mistakes: substitutions (wrong word), insertions (extra word), and deletions (missing word). Add them up, divide by the number of words actually spoken:

WER = (substitutions + insertions + deletions) ÷ words spoken. Accuracy, as vendors quote it, is simply 100% minus WER.

A toy example. Someone says: *"send the final report to the client by Friday morning"* — ten words. The engine writes: *"sent the a final report to the client by Friday."*

ErrorWhat happenedType
"send" → "sent"Wrong word for the word spokenSubstitution
"a" appearsA word nobody saidInsertion
"morning" is goneA spoken word never transcribedDeletion

Three errors over ten words: WER = 3 ÷ 10 = 30%, so this transcript is "70% accurate." Notice the deletion is arguably the worst mistake in the sentence — the deadline lost its time — and the insertion is nearly harmless. WER counts them the same. Hold that thought.

What state-of-the-art streaming accuracy feels like in an hour

state-of-the-art streaming accuracy means a 1.5% word error rate: roughly 1.5 words wrong in every 100. Percentages hide scale, so do the arithmetic at real volume. Continuous conversational speech runs to thousands of words an hour — if an hour of talk contains around 8,000 words, a 1.5% error rate is on the order of 120 imperfect words across it, a couple every minute of solid speech.

Whether that is excellent or unacceptable depends entirely on the job. For following a meeting you could mostly hear anyway, it is superb — the surrounding context absorbs nearly every slip. For a contract read aloud verbatim, no live system should be your only line of defense. For an interpreter using captions as a numbers-and-names safety net, it is transformative, because the words that matter most — figures, dosages, proper nouns — are exactly the ones worth engineering for with a custom dictionary.

Calibration: every 0.5 points of accuracy at this end of the scale is a big deal. Going from 97% to 98.5% doesn't sound like much — it is a 50% reduction in errors, from 3 per hundred words down to 1.5. Read accuracy deltas as error-rate ratios and vendor claims get much easier to rank.

Why the same engine scores differently on different audio

There is no such thing as "the" WER of an engine. There is only the WER of an engine on a particular set of recordings. Move the audio, move the number:

  • Audio quality. A close microphone in a quiet room versus a speakerphone across a conference table can be the difference between a clean transcript and a mess — same engine, same speaker, same words.
  • Accents and speaking style. Engines are strongest on the speech they saw most in training. Accented speech, fast talkers, mumbling, and code-switching all push the number up — by how much varies engine to engine.
  • Domain terminology. Drug names, statutes, product SKUs, surnames. An engine that has never seen "apixaban" will confidently write something else.
  • Overlap and crosstalk. Two people talking at once is the hardest ordinary condition there is; WER measured on polite turn-taking says little about a heated meeting.

This is also why vendor numbers and your experience can both be true. Published figures are measured on benchmark audio — typically cleaner, better recorded, and more standard-accented than your Tuesday afternoon call with two speakerphones and a sales team's product vocabulary. The benchmark isn't a lie; it just isn't your audio. Many of the conditions are under your control, and fixing them is the cheapest accuracy upgrade available — see how to improve live caption accuracy.

Not all errors are equal — and WER can't tell

WER is a word counter, not a meaning counter. It scores "no" transcribed as "now" — a negation silently inverted — exactly the same as a dropped "um." In practice the errors that hurt cluster in a few categories:

Negation and modality flips — "can" vs "can't", "no" vs "now". Rare, but each one can invert a sentence's meaning.
Numbers — dosages, amounts, dates, case numbers. A single wrong digit is a 1-word error with outsized consequences.
Proper nouns — names of people, drugs, companies, places. High-stakes and highly guessable-wrong.
Harmless noise — filler words, false starts, "uh" and "um". These inflate WER while costing nothing; an engine that drops them may read better than its score.

So when you evaluate, don't just count errors — read them. A tool with slightly higher WER but errors concentrated in harmless filler can beat a lower-WER tool that fumbles numbers. And remember from how streaming recognition works: judge the final text, not the interim flicker that corrects itself.

Run your own five-minute accuracy test

Pick five minutes of your real audio

Your accents, your vocabulary, your microphones — a slice of an actual meeting or call, not a podcast. The whole point is to measure the audio you will actually feed the tool.

Caption it live and save what you see

Run the tool under test on that audio in real conditions and capture its final output.

Make a reference transcript

Listen carefully and write down what was truly said, or hand-correct the machine output word by word. This is the tedious part; five minutes of audio keeps it bearable.

Count S + I + D and read the errors

Tally substitutions, insertions, and deletions against the reference; divide by reference word count for your personal WER. Then read the error list: how many were numbers, names, or meaning flips? That reading matters more than the percentage.

Frequently asked

What is word error rate and how is it calculated?

Word error rate (WER) measures transcription accuracy by comparing machine output to what was actually said: add up substitutions (wrong words), insertions (extra words), and deletions (missing words), then divide by the number of words spoken. Three errors in a ten-word sentence is a 30% WER, or "70% accuracy." Vendor accuracy claims are simply 100% minus WER.

What does 98.5% transcription accuracy actually mean?

It means a 1.5% word error rate — about 1.5 words wrong per 100 spoken, which over an hour of continuous speech adds up to a couple of imperfect words per minute. Context absorbs most of them; the ones that matter are numbers, names, and meaning flips like "can" vs "can't." The honest advice for any vendor's published accuracy number is the same: verify it on your own audio — the free tier gives you 30 minutes every week.

Why is my transcription accuracy worse than the vendor's claim?

Because WER is a property of the engine and the audio together, and published figures are measured on benchmark recordings — cleaner, closer-miked, and more standard-accented than most real calls. Speakerphones, crosstalk, accents, and domain jargon all raise the error rate. Improve the audio path and preload terminology in a custom dictionary before concluding the engine is at fault; both routinely move the number more than switching vendors does.

How can I test a captioning tool's accuracy myself?

Take five minutes of your own real audio, run the tool live on it, then make a careful reference transcript of what was truly said. Count substitutions, insertions, and deletions against the reference and divide by the reference word count — that is your personal WER. Then read the individual errors: a tool whose mistakes are filler words beats one that fumbles numbers, even at the same score. Unicaption's free free weekly 30 minutes exist for exactly this test.

Measure it on your audio

Any accuracy claim — verified by you, on your meetings, in five minutes. 30 free minutes every week, no credit card.

Start Free Trial →