Communication & Media Codexery

Speech processing

Study of speech signals and their digital processing methods.

Speech processing

Speech processing is the study of speech signals and the processing methods of those signals. The signals are usually processed in a digital representation, making speech processing a special case of digital signal processing applied to speech signals. Aspects of speech processing include the acquisition, manipulation, storage, transfer, and output of speech signals, with tasks such as speech recognition, speech synthesis, speaker diarization, speech enhancement, and speaker recognition.

field
Digital signal processing, speech recognition, speech synthesis
known_for
Enabling voice-operated systems, virtual assistants, and call center automation through techniques like linear predictive coding, hidden Markov models, and deep neural networks

Lore & Background

Early attempts at speech processing and recognition were primarily focused on understanding a handful of simple phonetic elements such as vowels. Biddulph, and K. H. Davis—developed a system that could recognize digits spoken by a single speaker. Pioneering works in the field of speech recognition using analysis of its spectrum were reported in the 1940s. Further developments in LPC technology were made by Bishnu S. Atal and Manfred R. Schroeder at Bell Labs during the 1970s.

Reader's Guide

Speech processing has evolved from early vowel recognition to sophisticated systems that power modern virtual assistants. The development of linear predictive coding in the 1960s and 1970s laid the groundwork for voice-over-IP and speech synthesis. By the early 2000s, the dominant strategy shifted from hidden Markov models to neural networks and deep learning. In 2012, Geoffrey Hinton and his team at the University of Toronto demonstrated that deep neural networks could significantly outperform traditional HMM-based systems on large vocabulary continuous speech recognition tasks, leading to widespread industry adoption. By the mid-2010s, companies like Google, Microsoft, Amazon, and Apple integrated advanced speech recognition into virtual assistants such as Google Assistant, Cortana, Alexa, and Siri. Transformer-based models like BERT and GPT further pushed boundaries, enabling more context-aware understanding. End-to-end speech recognition models have recently gained popularity by directly converting audio input into text output, streamlining development and improving performance.

More in Communication & Media 1-19

Spotted an error? Know more?

This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record

Comments

Loading…
Open in the interactive codex →