کد مقاله | کد نشریه | سال انتشار | مقاله انگلیسی | نسخه تمام متن |
---|---|---|---|---|
558276 | 874889 | 2014 | 22 صفحه PDF | دانلود رایگان |
• Review the state-of-the-art methods in glottal source processing.
• Tackle synchronization issues: pitch tracking, speech polarity and GCI detection.
• Review the techniques for glottal source estimation and parameterization.
• Emphasize the current voice technology applications based on the glottal source.
• Provide clues in identifying challenges in glottal source processing.
The great majority of current voice technology applications rely on acoustic features, such as the widely used MFCC or LP parameters, which characterize the vocal tract response. Nonetheless, the major source of excitation, namely the glottal flow, is expected to convey useful complementary information. The glottal flow is the airflow passing through the vocal folds at the glottis. Unfortunately, glottal flow analysis from speech recordings requires specific and complex processing operations, which explains why it has been generally avoided. This paper gives a comprehensive overview of techniques for glottal source processing. Starting from analysis tools for pitch tracking, detection of glottal closure instant, estimation and modeling of glottal flow, this paper discusses how these tools and techniques might be properly integrated in various voice technology applications.
Journal: Computer Speech & Language - Volume 28, Issue 5, September 2014, Pages 1117–1138