Patent search ap:("GOOGLE LLC") AND inv:"Ignacio Lopez Moreno" Page 11

101.

发明授权
Training and/or using a language selection model for automatically determining language for speech recognition of spoken utterance 有权

公开(公告)号：US11646011B2

公开(公告)日：2023-05-09

申请号：US17846287

申请日：2022-06-22

Applicant: Google LLC

Inventor： Li Wan , Yang Yu , Prashant Sridhar , Ignacio Lopez Moreno , Quan Wang

IPC: G10L15/00

CPC classification number: G10L15/005

Abstract: Methods and systems for training and/or using a language selection model for use in determining a particular language of a spoken utterance captured in audio data. Features of the audio data can be processed using the trained language selection model to generate a predicted probability for each of N different languages, and a particular language selected based on the generated probabilities. Speech recognition results for the particular language can be utilized responsive to selecting the particular language of the spoken utterance. Many implementations are directed to training the language selection model utilizing tuple losses in lieu of traditional cross-entropy losses. Training the language selection model utilizing the tuple losses can result in more efficient training and/or can result in a more accurate and/or robust model—thereby mitigating erroneous language selections for spoken utterances.

102.

发明授权
Speaker verification 有权

公开(公告)号：US11594230B2

公开(公告)日：2023-02-28

申请号：US17307704

申请日：2021-05-04

Applicant: Google LLC

Inventor： Ignacio Lopez Moreno , Li Wan , Quan Wang

IPC: G10L17/24 , G10L17/22 , G10L17/02 , G10L17/08 , G10L17/14 , G10L17/18

Abstract: Methods, systems, apparatus, including computer programs encoded on computer storage medium, to facilitate language independent-speaker verification. In one aspect, a method includes actions of receiving, by a user device, audio data representing an utterance of a user. Other actions may include providing, to a neural network stored on the user device, input data derived from the audio data and a language identifier. The neural network may be trained using speech data representing speech in different languages or dialects. The method may include additional actions of generating, based on output of the neural network, a speaker representation and determining, based on the speaker representation and a second representation, that the utterance is an utterance of the user. The method may provide the user with access to the user device based on determining that the utterance is an utterance of the user.

103.

发明申请
Speaker Identification Accuracy 有权

公开(公告)号：US20230015169A1

公开(公告)日：2023-01-19

申请号：US17933164

申请日：2022-09-19

Applicant: Google LLC

Inventor： Yeming Fang , Quan Wang , Pedro Jose Moreno Mengibar , Ignacio Lopez Moreno , Gang Feng , Fang Chu , Jin Shi , Jason William Pelecanos

IPC: G10L17/06

Abstract: A method of generating an accurate speaker representation for an audio sample includes receiving a first audio sample from a first speaker and a second audio sample from a second speaker. The method includes dividing a respective audio sample into a plurality of audio slices. The method also includes, based on the plurality of slices, generating a set of candidate acoustic embeddings where each candidate acoustic embedding includes a vector representation of acoustic features. The method further includes removing a subset of the candidate acoustic embeddings from the set of candidate acoustic embeddings. The method additionally includes generating an aggregate acoustic embedding from the remaining candidate acoustic embeddings in the set of candidate acoustic embeddings after removing the subset of the candidate acoustic embeddings.

104.

发明授权
Speaker diartzation using an end-to-end model 有权

公开(公告)号：US11545157B2

公开(公告)日：2023-01-03

申请号：US16617219

申请日：2019-04-15

Applicant: Google LLC

Inventor： Quan Wang , Yash Sheth , Ignacio Lopez Moreno , Li Wan

IPC: G10L17/18 , G10L15/26 , G10L17/04 , G10L21/0216 , G06K9/62 , G10L15/16 , G10L17/00

Abstract: Techniques are described for training and/or utilizing an end-to-end speaker diarization model. In various implementations, the model is a recurrent neural network (RNN) model, such as an RNN model that includes at least one memory layer, such as a long short-term memory (LSTM) layer. Audio features of audio data can be applied as input to an end-to-end speaker diarization model trained according to implementations disclosed herein, and the model utilized to process the audio features to generate, as direct output over the model, speaker diarization results. Further, the end-to-end speaker diarization model can be a sequence-to-sequence model, where the sequence can have variable length. Accordingly, the model can be utilized to generate speaker diarization results for any of various length audio segments.

105.

发明申请
SPEAKER AWARENESS USING SPEAKER DEPENDENT SPEECH MODEL(S) 有权

公开(公告)号：US20220157298A1

公开(公告)日：2022-05-19

申请号：US17587424

申请日：2022-01-28

Applicant: GOOGLE LLC

Inventor： Ignacio Lopez Moreno , Quan Wang , Jason Pelecanos , Li Wan , Alexander Gruenstein , Hakan Erdogan

IPC: G10L15/06 , G10L15/07 , G10L15/20 , G10L17/04 , G10L17/20 , G10L21/0208

Abstract: Techniques disclosed herein enable training and/or utilizing speaker dependent (SD) speech models which are personalizable to any user of a client device. Various implementations include personalizing a SD speech model for a target user by processing, using the SD speech model, a speaker embedding corresponding to the target user along with an instance of audio data. The SD speech model can be personalized for an additional target user by processing, using the SD speech model, an additional speaker embedding, corresponding to the additional target user, along with another instance of audio data. Additional or alternative implementations include training the SD speech model based on a speaker independent speech model using teacher student learning.

106.

发明申请
SPEAKER AWARENESS USING SPEAKER DEPENDENT SPEECH MODEL(S) 有权

公开(公告)号：US20210312907A1

公开(公告)日：2021-10-07

申请号：US17251163

申请日：2019-12-04

Applicant: GOOGLE LLC

Inventor： Ignacio Lopez Moreno , Quan Wang , Jason Pelecanos , Li Wan , Alexander Gruenstein , Hakan Erdogan

IPC: G10L15/06 , G10L15/07 , G10L21/0208 , G10L15/20 , G10L17/04 , G10L17/20

Abstract: Techniques disclosed herein enable training and/or utilizing speaker dependent (SD) speech models which are personalizable to any user of a client device. Various implementations include personalizing a SD speech model for a target user by processing, using the SD speech model, a speaker embedding corresponding to the target user along with an instance of audio data. The SD speech model can be personalized for an additional target user by processing, using the SD speech model, an additional speaker embedding, corresponding to the additional target user, along with another instance of audio data. Additional or alternative implementations include training the SD speech model based on a speaker independent speech model using teacher student learning.

107.

发明申请
AUTOMATICALLY DETERMINING LANGUAGE FOR SPEECH RECOGNITION OF SPOKEN UTTERANCE RECEIVED VIA AN AUTOMATED ASSISTANT INTERFACE 有权

公开(公告)号：US20210280177A1

公开(公告)日：2021-09-09

申请号：US17328400

申请日：2021-05-24

Applicant: Google LLC

Inventor： Pu-sen Chao , Diego Melendo Casado , Ignacio Lopez Moreno , William Zhang

IPC: G10L15/197 , G10L15/00 , G10L15/22 , G10L15/30 , G10L15/08 , G10L15/14 , G10L15/18 , G10L13/00

Abstract: Determining a language for speech recognition of a spoken utterance received via an automated assistant interface for interacting with an automated assistant. Implementations can enable multilingual interaction with the automated assistant, without necessitating a user explicitly designate a language to be utilized for each interaction. Implementations determine a user profile that corresponds to audio data that captures a spoken utterance, and utilize language(s), and optionally corresponding probabilities, assigned to the user profile in determining a language for speech recognition of the spoken utterance. Some implementations select only a subset of languages, assigned to the user profile, to utilize in speech recognition of a given spoken utterance of the user. Some implementations perform speech recognition in each of multiple languages assigned to the user profile, and utilize criteria to select only one of the speech recognitions as appropriate for generating and providing content that is responsive to the spoken utterance.

108.

发明授权
Neural networks for speaker verification 有权

公开(公告)号：US11107478B2

公开(公告)日：2021-08-31

申请号：US16752007

申请日：2020-01-24

Applicant: Google LLC

Inventor： Georg Heigold , Samuel Bengio , Ignacio Lopez Moreno

IPC: G10L17/18 , G10L17/04 , G10L17/02

Abstract: This document generally describes systems, methods, devices, and other techniques related to speaker verification, including (i) training a neural network for a speaker verification model, (ii) enrolling users at a client device, and (iii) verifying identities of users based on characteristics of the users' voices. Some implementations include a computer-implemented method. The method can include receiving, at a computing device, data that characterizes an utterance of a user of the computing device. A speaker representation can be generated, at the computing device, for the utterance using a neural network on the computing device. The neural network can be trained based on a plurality of training samples that each: (i) include data that characterizes a first utterance and data that characterizes one or more second utterances, and (ii) are labeled as a matching speakers sample or a non-matching speakers sample.

109.

发明授权
Speaker diarization using speaker embedding(s) and trained generative model 有权

公开(公告)号：US10978059B2

公开(公告)日：2021-04-13

申请号：US16607977

申请日：2018-09-25

Applicant: Google LLC

Inventor： Ignacio Lopez Moreno , Luis Carlos Cobo Rus

IPC: G10L15/00 , G10L15/20 , G10L15/30 , G10L15/02 , G10L21/0208 , G10L15/06

Abstract: Speaker diarization techniques that enable processing of audio data to generate one or more refined versions of the audio data, where each of the refined versions of the audio data isolates one or more utterances of a single respective human speaker. Various implementations generate a refined version of audio data that isolates utterance(s) of a single human speaker by generating a speaker embedding for the single human speaker, and processing the audio data using a trained generative model—and using the speaker embedding in determining activations for hidden layers of the trained generative model during the processing. Output is generated over the trained generative model based on the processing, and the output is the refined version of the audio data.

110.

发明授权
Speech recognition using neural networks 有权

公开(公告)号：US10930271B2

公开(公告)日：2021-02-23

申请号：US16573232

申请日：2019-09-17

Applicant: Google LLC

Inventor： Andrew W. Senior , Ignacio Lopez Moreno

IPC: G10L15/16 , G06N3/02 , G10L15/02

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech recognition using neural networks. A feature vector that models audio characteristics of a portion of an utterance is received. Data indicative of latent variables of multivariate factor analysis is received. The feature vector and the data indicative of the latent variables is provided as input to a neural network. A candidate transcription for the utterance is determined based on at least an output of the neural network.

Search Results

Country/Region

Patent validity

Application date

Publication (announcement) day

applicant

The country/region where the applicant is located

Inventor

IPC

IPC Department

IPC class

IPC subclass

IPC group

IPC team

Appearance classification