Glossary
The terms, briefly explained
The field is full of abbreviations and words that mean something else in everyday speech. Here are 30 of them, explained plainly, short enough to finish.
Speech recognition and models
- ASR
- Automatic speech recognition, the technical term for turning speech into text automatically.
- Code-switching
- Shifting between two or more languages in the same conversation, often within a single sentence.
- Diarisation
- Splitting a recording into segments by who is speaking, without necessarily knowing who those people are.
- Fine-tuning
- Training a finished model further on a narrower material, for instance speech in one language.
- Hallucination
- Text the model produces without there being any basis for it in the audio.
- Language model
- A model that predicts likely word sequences, used both in transcription and to write summaries.
- NB-Whisper
- The National Library of Norway's further-trained version of Whisper, adapted to Norwegian speech and dialect.
- Speaker recognition
- Tying a voice to a known person, not merely telling it apart from other voices.
- Speech recognition
- The technology that turns the sound of speech into written text.
- Voice activity detection (VAD)
- The technique that decides where in an audio recording somebody is speaking, and where it is silent.
- Whisper
- An open speech recognition model from OpenAI that can be run on your own machine.
- Word error rate (WER)
- The share of words that are wrong, missing or added, divided by the number of words in the reference.
Methods and workflow
- Batch transcription
- Transcribing a finished recording afterwards, rather than while the speaking happens.
- Cloud transcription
- Transcription where the audio is uploaded and processed on the vendor's servers.
- Local transcription
- Transcription that runs on your own machine, without the audio being sent to a server.
- Post-editing
- Going through and correcting a machine-generated transcript against the audio.
- Real-time transcription
- Transcription that happens while the speaking happens, with one to three seconds of delay.
File formats and audio
- Audio format
- The file format the audio is stored in, for instance m4a, mp3 or wav.
- Captioning
- Showing speech as readable text on screen, timed and limited in lines and characters.
- Sample rate
- How many times per second the audio signal is measured, given in hertz.
- SRT
- The most common subtitle format: numbered text blocks with a start and end time.
- Timestamp
- A time reference tying a piece of the text to a particular point in the audio recording.
- VTT
- The web standard for subtitles, with support for styling, positioning and metadata.
Transcript forms
- Edited transcript
- A transcript where the content is kept, but phrased in whole, readable sentences.
- Verbatim
- Word-for-word transcription that includes hesitation, repetition, interruptions and often pauses.
Meetings and minutes
- Action item
- A concrete task following from a meeting, with one owner and one deadline.
- Any other business
- The standing last item on the agenda, for short matters not submitted in advance.
- Decision log
- A running list of decisions across meetings, with date, owner and reasoning.
- Dissent
- A member voting against, or recording disagreement with, a decision.
- Statement for the record
- A declaration a member requires to be entered in the minutes, regardless of what the majority thinks.
Why English dominates the field
Research on speech recognition was done in English, and the terms came along with it. Some languages have good native words, such as the Norwegian ordfeilrate for word error rate. Others do not, and then the English term beats inventing something nobody recognises. Where both are in use, both are listed here.
Speech, written out
Sayable transcribes speech with dialect, tells the speakers apart and drafts the summary. Free to get started.