?
Method of voice source coding with data compression based on the linear prediction model
The problem of voice source coding with data compression based on the linear prediction model is considered
within the rapidly developing research area in the field of acoustic measurements—parameter analysis and
estimation of an excitation signal inducing acoustic oscillations in a speaker’s vocal tract. With the application
of the criterion of minimum average voice source power in speech production, the described problem is reduced
to the real-time coding of a linear prediction error signal. A voice coding method that involves clipping of
the linear prediction error was developed. The proposed method provides a means to avoid computationally
intensive procedures for measuring the initial phase and fundamental frequency of a speech signal. An example
of its technical implementation in soft real-time mode is considered. For a comparative effectiveness analysis
of the proposed method and the widespread method of discrete cosine transform (DCT), a full-scale experiment
was set up and conducted. It is shown that due to the reduction of data compression artifacts in the reconstructed
speech signal, the accuracy of voice source coding via the developed method is one and a half to two times
higher (as compared to the DCT method), and it is not necessary to detect vowel speech sounds and pauses in
the speech signal. The obtained results can be used to develop new and upgrade existing systems and algorithms
in the field of automatic speech processing and synthesis, mobile speech communication, artificial intelligence,
and other applications of speech technologies with data compression based on the linear prediction model.