Rozpoznawanie mowy metodami niejawnych modeli Markowa HMM
Relation
Local access
Defence Date
2007
Degree Date
Authors
Supervisors:
Reviewers:
Other title
Recognition of speech through hidden Markov models HMM
Resource type
Call number
Defence details
Physical Description:
Research Project
Description
Abstract
The subject of the dissertation is the application of hidden Markov models for recognition of speech. In the dissertation there were discussed basic optimization algorithms for models DHMM and CDHMM, i.e. Baum-Welch algorithm, segmental k>means algorithm and Bayes adaptation algorithm. In the dissertation the network model DHMM formulated by the author as well as methods of its optimization have been discussed. The essential advantage of the discussed network model is that it may be used for automatic language segmentation. For DHMM models there were placed statements of calculation and memory complexity of the algorithm of determination of observation probability and Baum-Welch algorithm. Since there do not exist methods for determination of global optimum for probability function P(O|M), therefore for optimization of DHMM models there were used algorithms of artificial intelligence, i.e. genetic algorithm, and ILS algorithm (Iterated Local Search). Within the dissertation the program of speech recognition based on the discussed DHMM models was created. In the program two methods of transformation of sound signal were applied, i.e. LPC and MFCC.
Przedmiotem rozprawy jest zastosowanie niejawnych modeli Markowa do rozpoznawania mowy. W rozprawie zostały omówione podstawowe algorytmy optymalizacji dla modeli DHMM oraz CDHMM tj: Algorytm Bauma-Welcha, algorytm segmental k-means oraz algorytm adaptacji Bayesa. W rozprawie został omówiony opracowany przez autora model sieciowy DHMM oraz metody jego optymalizacji. Istotną zaletą omówionego modelu sieciowego jest to, że może on zostać wykorzystany do automatycznego podziału językowego. Dla modeli DHMM zostało zamieszczone zestawienie złożoności obliczeniowej oraz pamięciowej algorytmu wyznaczania prawdopodobieństwa obserwacji oraz algorytmu Bauma-Welch. Ponieważ nie ma metod wyznaczania optimum globalnego dla funkcji prawdopodobieństwa P(O|M), dlatego do optymalizacji modeli DHMM zostały wykorzystane algorytmy sztucznej inteligencji tj: algorytm Genetyczny, oraz algorytm ILS (Iterated Local Search). W ramach rozprawy powstał program rozpoznawania mowy oparty na omówionych modelach DHMM. W programie zostały zastosowane dwie metody przetwarzania sygnału dźwięku tj. metody: LPC oraz MFCC.

