Glossary · Multimodal systems

Automatic Speech Recognition (ASR)

The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence information.

Why it matters

Speech interfaces depend on more than language modeling. Acoustic variation, segmentation, decoding, vocabulary, and domain conditions all affect the final transcript.

In practice

Evaluate word or character errors by language, speaker, noise, and domain, retain timestamps when downstream grounding needs them, and test the exact audio preprocessing used in production.

Common confusion

ASR transcribes what was said. Determining who spoke requires diarization or speaker recognition, while translation and intent understanding are separate tasks.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.