摘要 |
During the recognition, speech values which are derived from sample values of the speech signals are compared with reference values, the words of a given vocabulary each time being given by a sequence of reference values. The words are then determined from phonemes according to a fixed pronouncing lexicon and the reference values for the phonemes are determined in a learning phase, each phoneme within a word consisting of a number of equal reference values determined in the learning phase. In order to approach transitions between phonemes, each phoneme may also consist of three sections of each time constant reference values. By the given number of reference values per phoneme, the time duration of a phoneme in a given word can be simulated more accurately. Different possibilities are indicated to determine the reference values and the distance value during the recognition. |