0:00
/
Generate transcript
A transcript unlocks clips, previews, and editing.

Xiaochang Li: Speech Recognition and the Limits of Imitation

Aryan Talks Tech #009

In the history of computing, there have been many people trying to build machines that could recognize speech. For much of the twentieth century, they agreed that the machine had to produce something inherently human. At Bell Labs in the 1940s and ‘50s, that meant building an “artificial ear,” a device modeled on how the inner ear processes sound. By the 1970s, the ARPA Speech Understanding Research program had made the ambition far larger: encode the knowledge of phoneticians and linguists into software that would not just recognize speech but understand it. Then a group at IBM led by the information theorist Fred Jelinek decided that asking a machine to emulate people was the mistake.

Xiaochang Li, Assistant Professor of Communication at Stanford, studies how language became something computers could process. We trace this history from the spectrograph to the noisy-channel model, and she explains why calling IBM’s break simply “the statistical approach” misses what actually changed.

Discussion about this video

User's avatar

Ready for more?