Speech recognition and models

Speech recognition

The technology that turns the sound of speech into written text.

·Also called: ASR, automatic speech recognition

Speech recognition is the technology that takes in the sound of speech and gives back text. It sits underneath everything from dictation on your phone to transcribing several hours of meeting audio.

The phrase is used for two slightly different things, and they are worth keeping apart:

Speech recognition as technology. The model that converts sound into letters. This is the ordinary meaning, and automatic speech recognition, or ASR, covers the same ground.

Speech recognition as identification. Recognising who is speaking, not what is being said. This is better called speaker recognition. The two get mixed up often, and the difference matters: one writes down words, the other recognises a person.

Not the same as understanding language

A speech recognition system writes down what was said. It does not understand what it means, and it does not know that “yes, that will probably be fine” can mean the opposite. Interpretation happens further down the line, in a language model or in the head of whoever reads it.

The distinction also explains why these systems miss proper nouns and jargon but handle ordinary sentences well: they are trained on which sound patterns and word sequences are likely, not on what makes sense inside your organisation.

See also

ASR, speaker recognition, word error rate.

Speech, written out

Sayable transcribes speech with dialect, tells the speakers apart and drafts the summary. Free to get started.