Abstract:
A data processing system holds metadata extraction dictionary information defining a condition for extracting metadata from a plurality of kinds of data and relevance dictionary information defining a condition for associating the metadata extracted from the plurality of kinds of data, extracts the metadata from the plurality of kinds of data on the basis of the metadata extraction dictionary information, extracts metadata from inputted data, associates the metadata extracted from the inputted data with the metadata extracted from the plurality of kinds of data on the basis of the relevance dictionary information, and outputs information indicating a relation of any combination of the plurality of kinds of data, the inputted data, and the metadata extracted from the plurality of kinds of data and the inputted data on the basis of a result of the association.
Abstract:
An audio analysis platform may receive a portion of an audio input, wherein the audio input corresponds to audio associated with a plurality of speakers. The audio analysis platform may process, using a neural network, the portion of the audio input to determine voice activity of the plurality of speakers during the portion of the audio input, wherein the neural network is trained using reference audio data and reference diarization data corresponding to the reference audio data. The audio analysis platform may determine, based on the neural network being used to process the portion of the audio input, a diarization output associated with the portion of the audio input, wherein the diarization output indicates individual voice activity of the plurality of speakers. The audio analysis platform may provide the diarization output to indicate the individual voice activity of the plurality of speakers during the portion of the audio input.
Abstract:
Provided is a voice search technology that can efficiently find and check a problematic call. To this end, a voice search system of the present invention includes a call search database that stores, for each of a reception channel and a transmission channel of each of a plurality of pieces of recorded call voice data, voice section sequences in association with predetermined keywords and time information. The call search database is searched based on an input search keyword, so that a voice section sequence that contains the search keyword is obtained. More specifically, the voice search system obtains, as a keyword search result, a voice section sequence that contains the search keyword and the appearance time thereof from the plurality of pieces of recorded call voice data, and obtains, based on the appearance time in the keyword search result, the start time of a voice section sequence of another channel immediately before the voice section sequence obtained as the keyword search result, and thus determines the start time as the playback start position for playing back the recorded voice. Then, the playback start position is output as a voice search result.