Glossary · Multimodal systems
Automatic Speech Recognition (ASR)
The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence information.
Why it matters
Speech interfaces depend on more than language modeling. Acoustic variation, segmentation, decoding, vocabulary, and domain conditions all affect the final transcript.
In practice
Evaluate word or character errors by language, speaker, noise, and domain, retain timestamps when downstream grounding needs them, and test the exact audio preprocessing used in production.
Common confusion
ASR transcribes what was said. Determining who spoke requires diarization or speaker recognition, while translation and intent understanding are separate tasks.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.