International Workshop on Speech Dynamics by Ear, Eye, Mouth and Machine, Technical Report of IEICE Japan, Vol. SP2003-48, pp. 27-36, 2003 (Invited Paper)
What are the essential cues for understanding spoken language?
S. Greenberg and T. Arai
Abstract: Classical models of speech recognition assume that a detailed, short-term analysis of the acoustic signal is essential for accurately decoding the speech signal and that this deconding process is rooted in the phonetic segment. This paper presents an altemative view, one in which the time scales required to accurately describe and model spoken language are both shorter and longer than the phonetic segment, and are inherently wedded to the syllable. The syllable reflects a singular property of the acoustic signal – the modulation spectrum – which provides a principled, quantitative framework to describe the process by which the listener proceeds from sound to meaning. The ability to understand spoken language (i.e.,intelligibility) vitally depends on the integrity of the modulation spectrum within the core range of the syllable (3-10 Hz) and reflects the variation in syllable emphasis associated with the concept of prosodie prominence (“accent”). A model of spoken language is deseribed in which the properties of the speech signal are embedded in the temporal dynamics associated with the syllable, a unit serving as the organizational interface among the various tiers of linguistic representation.
Keywords: Speech Perception, Intelligibility, Syllables, Modulation Spectrum, Cross-Spectral Asynchrony, Auditory System