摘要:
The system implements high-accuracy speech recognition while suppressing the amount of data transfer between the client and server. For this purpose, the client compression-encodes speech parameters by a speech processing unit, and sends the compression-encoded speech parameters to the server. The server receives the compression-encoded speech parameters, a speech processing unit makes speech recognition of the compression-encoded speech parameters, and sends information corresponding to the speech recognition result to the client.
摘要:
A GUI display module displays a contents image based on contents data within a display area, and a display portion switching input module instructs to change the display portion of the contents image within the display area. Based on this instruction input, a display portion switching module changes the display portion of the contents image within the display area. A synthesis text determination module determines data which is to undergo speech synthesis in the contents data on the basis of display portion information which is held by a display portion holding module and indicates the display portion. A speech synthesis module synthesizes speech of the data which is to undergo speech synthesis, and a speech output module outputs the synthesized synthetic speech.
摘要:
A method and apparatus for recognizing speech employing a word dictionary in which the phoneme of words are stored and for recognizing speech based on the recognition of the phonemes. The method and apparatus recognize phonemes and produce data associated with each phoneme according to different speech analyzing and recognizing methods for each kind of phoneme, normalize the produced data, and match the recognized phonemes with words in the word dictionary by means of dynamic programming based on the normalized data.
摘要:
A method for encoding syllables of a language, particularly the Japanese language, and for facilitating the extraction of sound codes from the input syllables, for voice recognition or voice synthesis includes the step of providing a syllable classifying table, in which each syllable is represented by an upper byte code indicating the consonant part of the syllable and a lower byte code indicating the non-consonant part of the syllable. The consonants constitute a first category of data classified by phonetic features, while the non-consonants constitute a second category of data classified by phonetic features, so that the extraction of consonant or non-consonant sounds can be made by a search in only the first or the second categories. The encoding of diphthongs are made in such a manner that those containing the same vowel have the same remainder corresponding to the code of this vowel, when the codes are divided by the number of vowels contained in the second category, so that the extraction of a vowel from diphthongs can be achieved by a simple mathematical division.
摘要:
A user dictionary, which is formed by storing pronunciations and notations of target recognition words designated by the user in correspondence with each other, input speech recognition data, and dictionary management data used to determine the recognition field of a recognition dictionary used in recognition of the speech recognition data are sent to a server via a communication module. In the server, a dictionary management unit looks up an identifier table to determine a recognition dictionary corresponding to the dictionary management information received from a client from a plurality of kinds of recognition dictionaries. A speech recognition module recognizes the speech recognition data using at least the determined recognition dictionary. The recognition result is sent to the client via a communication module.
摘要:
The present invention is characterized in that: a demanded printing-job is accepted at a printing-job accepting unit, and a printing executing unit executes printing according to the accepted printing-job so as to produce an output sheet onto a discharge tray; an output-sheet detecting unit detects the presence of a printed output-sheet on the discharge tray; if the removal of discharged sheets from the discharge tray is detected, a printed-sheet mixture determining unit determines the possibility that sheets printed corresponding to a plurality of demands are mixed among the removed output-sheets; and if it is determined the possibility that sheets printed corresponding to a plurality of the demands are mixed among the output-sheets, warning information is produced by a warning voice producing unit.
摘要:
The system implements high-accuracy speech recognition while suppressing the amount of data transfer between the client and server. For this purpose, the client compression-encodes speech parameters by a speech processing unit, and sends the compression-encoded speech parameters to the server. The server receives the compression-encoded speech parameters, and speech processing unit makes speech recognition of the compression-encoded speech parameters, and sends information corresponding to the speech recognition result to the client.
摘要:
An information processing apparatus inputs a document having a plurality of input items, and displays it using an information display unit. An active input field is discriminated from the plurality of input fields in accordance with the display state of the document. A specific grammar corresponding to the discriminated active input field is selected from a grammar holding unit for holding a plurality of types of grammars, and the selected grammar is used in a speech recognition process.
摘要:
An apparatus, method, and storage medium for eliminating the influence of line characteristics in a real-time manner in order to raise recognition precision of input speech and to enable the speech to be recognized in a real-time manner, includes a device and step for obtaining, an estimate value of a long-time mean of a parameter from speech feature parameters which are sequentially inputted by using the speech feature parameters which have already been inputted, and a device and step for normalizing the speech feature parameter inputted at that time point by using the obtained estimate value. Each time the speech feature parameter is inputted, the latest estimate value is obtained by using the already inputted parameters including the inputted speech feature parameter, and the latest input speech feature parameter is normalized by using the updated estimate value. Since the reliability of the estimate value is higher as the number of speech feature parameters used when the estimate value is obtained is larger, the estimate value is normalized by adding a weight in accordance with the reliability.
摘要:
Speech recognition is achieved using a normalized cumulative distance. A normalized Dynamic Programming (DP) value is calculated by dividing a cumulative path distance by an optimal integral path length. The path length is calculated iteratively by adding 2 if the warping path is diagonal or by adding 3 if the warping path is horizontal or vertical. Distance may be calculated by measuring a difference between input power and average power. The power difference is weighted by a coefficient (.lambda.) between 0 and 1. A Mahalanobis distance is then weighted by (1-.lambda.) and added to the weighted power difference.