Every accuracy percentage you have ever seen on a transcription product is a WER measurement in disguise. Here is how the math works, and how to read the claims like someone who knows.
"highly accurate." "Industry-leading accuracy." "Near-perfect transcription." Every one of these claims comes from the same underlying measurement — word error rate — and once you know how it is computed, you can tell which claims mean something and which are marketing weather. Accuracy is not a property of an engine; it is a property of an engine plus the audio you feed it — which is why the only number that should decide your purchase is the one you measure on your own recordings. The math takes two minutes to learn.
WER compares what the engine wrote against what was actually said, counting three kinds of mistakes: substitutions (wrong word), insertions (extra word), and deletions (missing word). Add them up, divide by the number of words actually spoken:
WER = (substitutions + insertions + deletions) ÷ words spoken. Accuracy, as vendors quote it, is simply 100% minus WER.
A toy example. Someone says: *"send the final report to the client by Friday morning"* — ten words. The engine writes: *"sent the a final report to the client by Friday."*
| Error | What happened | Type |
|---|---|---|
| "send" → "sent" | Wrong word for the word spoken | Substitution |
| "a" appears | A word nobody said | Insertion |
| "morning" is gone | A spoken word never transcribed | Deletion |
Three errors over ten words: WER = 3 ÷ 10 = 30%, so this transcript is "70% accurate." Notice the deletion is arguably the worst mistake in the sentence — the deadline lost its time — and the insertion is nearly harmless. WER counts them the same. Hold that thought.
state-of-the-art streaming accuracy means a 1.5% word error rate: roughly 1.5 words wrong in every 100. Percentages hide scale, so do the arithmetic at real volume. Continuous conversational speech runs to thousands of words an hour — if an hour of talk contains around 8,000 words, a 1.5% error rate is on the order of 120 imperfect words across it, a couple every minute of solid speech.
Whether that is excellent or unacceptable depends entirely on the job. For following a meeting you could mostly hear anyway, it is superb — the surrounding context absorbs nearly every slip. For a contract read aloud verbatim, no live system should be your only line of defense. For an interpreter using captions as a numbers-and-names safety net, it is transformative, because the words that matter most — figures, dosages, proper nouns — are exactly the ones worth engineering for with a custom dictionary.
There is no such thing as "the" WER of an engine. There is only the WER of an engine on a particular set of recordings. Move the audio, move the number:
This is also why vendor numbers and your experience can both be true. Published figures are measured on benchmark audio — typically cleaner, better recorded, and more standard-accented than your Tuesday afternoon call with two speakerphones and a sales team's product vocabulary. The benchmark isn't a lie; it just isn't your audio. Many of the conditions are under your control, and fixing them is the cheapest accuracy upgrade available — see how to improve live caption accuracy.
WER is a word counter, not a meaning counter. It scores "no" transcribed as "now" — a negation silently inverted — exactly the same as a dropped "um." In practice the errors that hurt cluster in a few categories:
So when you evaluate, don't just count errors — read them. A tool with slightly higher WER but errors concentrated in harmless filler can beat a lower-WER tool that fumbles numbers. And remember from how streaming recognition works: judge the final text, not the interim flicker that corrects itself.
Your accents, your vocabulary, your microphones — a slice of an actual meeting or call, not a podcast. The whole point is to measure the audio you will actually feed the tool.
Run the tool under test on that audio in real conditions and capture its final output.
Listen carefully and write down what was truly said, or hand-correct the machine output word by word. This is the tedious part; five minutes of audio keeps it bearable.
Tally substitutions, insertions, and deletions against the reference; divide by reference word count for your personal WER. Then read the error list: how many were numbers, names, or meaning flips? That reading matters more than the percentage.
Unicaption publishes state-of-the-art streaming accuracy — and everything above applies to us exactly as it applies to everyone else. It is a published figure; your audio is the real test, and we have made running that test free:
If you are comparing several tools, run the same five-minute test on each — our guide to the best live captioning tools of 2026 is a sensible shortlist to test against.
Word error rate (WER) measures transcription accuracy by comparing machine output to what was actually said: add up substitutions (wrong words), insertions (extra words), and deletions (missing words), then divide by the number of words spoken. Three errors in a ten-word sentence is a 30% WER, or "70% accuracy." Vendor accuracy claims are simply 100% minus WER.
It means a 1.5% word error rate — about 1.5 words wrong per 100 spoken, which over an hour of continuous speech adds up to a couple of imperfect words per minute. Context absorbs most of them; the ones that matter are numbers, names, and meaning flips like "can" vs "can't." The honest advice for any vendor's published accuracy number is the same: verify it on your own audio — the free tier gives you 30 minutes every week.
Because WER is a property of the engine and the audio together, and published figures are measured on benchmark recordings — cleaner, closer-miked, and more standard-accented than most real calls. Speakerphones, crosstalk, accents, and domain jargon all raise the error rate. Improve the audio path and preload terminology in a custom dictionary before concluding the engine is at fault; both routinely move the number more than switching vendors does.
Take five minutes of your own real audio, run the tool live on it, then make a careful reference transcript of what was truly said. Count substitutions, insertions, and deletions against the reference and divide by the reference word count — that is your personal WER. Then read the individual errors: a tool whose mistakes are filler words beats one that fumbles numbers, even at the same score. Unicaption's free free weekly 30 minutes exist for exactly this test.
Any accuracy claim — verified by you, on your meetings, in five minutes. 30 free minutes every week, no credit card.
Start Free Trial →