Speech recognition and models

Word error rate (WER)

The share of words that are wrong, missing or added, divided by the number of words in the reference.

·Also called: WER

Word error rate, abbreviated WER, is the standard measure of how accurate a transcription is.

The formula is simple:

WER = (substitutions + deletions + insertions) / number of words in the reference

All three error types count equally. A missing word weighs as much as a wrong one.

What the numbers mean

WERMeansIn practice
Under 5%19 out of 20 words rightLight editing, mostly proper nouns
5 to 10%9 out of 10 rightUsable, some editing
10 to 20%4 out of 5 rightHeavy going, but faster than typing
Over 20%Every fifth word wrongConsider typing it yourself

Why marketing figures are worth little

“98 percent accuracy” says nothing without stating what was measured. Almost every system gets above 95 percent on one person reading clearly from a script in a soundproofed room. No meeting looks like that.

The variables that actually decide it: how close the microphone was, how many people spoke, whether anyone spoke at once, which accent, and how much jargon. Without them the number is not comparable between vendors.

A limitation of the measure

WER weighs all words equally. A dropped “and” counts as much as a misspelled surname, even though the second is far worse for you. That is why it is worth counting proper nouns and jargon separately when you test.

See also

ASR, verbatim.

Speech, written out

Sayable transcribes speech with dialect, tells the speakers apart and drafts the summary. Free to get started.