Patent search ap:("Google LLC") AND inv:"Ignacio Lopez Moreno" Page 2

11.

发明授权
Speaker awareness using speaker dependent speech model(s) 有权

公开(公告)号：US11238847B2

公开(公告)日：2022-02-01

申请号：US17251163

申请日：2019-12-04

Applicant: GOOGLE LLC

Inventor： Ignacio Lopez Moreno , Quan Wang , Jason Pelecanos , Li Wan , Alexander Gruenstein , Hakan Erdogan

IPC: G10L17/00 , G10L15/06 , G10L15/07 , G10L15/20 , G10L17/04 , G10L17/20 , G10L21/0208 , G10L15/08

Abstract: Techniques disclosed herein enable training and/or utilizing speaker dependent (SD) speech models which are personalizable to any user of a client device. Various implementations include personalizing a SD speech model for a target user by processing, using the SD speech model, a speaker embedding corresponding to the target user along with an instance of audio data. The SD speech model can be personalized for an additional target user by processing, using the SD speech model, an additional speaker embedding, corresponding to the additional target user, along with another instance of audio data. Additional or alternative implementations include training the SD speech model based on a speaker independent speech model using teacher student learning.

12.

发明授权
Targeted voice separation by speaker conditioned on spectrogram masking 有权

公开(公告)号：US11217254B2

公开(公告)日：2022-01-04

申请号：US16598172

申请日：2019-10-10

Applicant: Google LLC

Inventor： Quan Wang , Prashant Sridhar , Ignacio Lopez Moreno , Hannah Muckenhirn

IPC: G10L17/04 , G10L17/22 , G10L25/18 , G10L17/02 , G10L17/18 , G10L17/00

Abstract: Techniques are disclosed that enable processing of audio data to generate one or more refined versions of audio data, where each of the refined versions of audio data isolate one or more utterances of a single respective human speaker. Various implementations generate a refined version of audio data that isolates utterance(s) of a single human speaker by processing a spectrogram representation of the audio data (generated by processing the audio data with a frequency transformation) using a mask generated by processing the spectrogram of the audio data and a speaker embedding for the single human speaker using a trained voice filter model. Output generated over the trained voice filter model is processed using an inverse of the frequency transformation to generate the refined audio data.

13.

发明申请
MULTI-USER AUTHENTICATION ON A DEVICE 有权

公开(公告)号：US20210343276A1

公开(公告)日：2021-11-04

申请号：US17375573

申请日：2021-07-14

Applicant: GOOGLE LLC

Inventor： Ignacio Lopez Moreno , Diego Melendo Casado

IPC: G10L15/08 , G06F21/32 , G10L17/06 , G06F16/635 , G10L15/22 , G10L17/00 , G06K9/00 , G10L15/07

Abstract: In some implementations, an utterance is determined to include a particular user speaking a hotword based at least on a first set of samples of the particular user speaking the hotword. In response to determining that an utterance includes a particular user speaking a hotword based at least on a first set of samples of the particular user speaking the hotword, at least a portion of the utterance is stored as a new sample. A second set of samples of the particular user speaking the utterance is obtained, where the second set of samples includes the new sample and less than all the samples in the first set of samples. A second utterance is determined to include the particular user speaking the hotword based at least on the second set of samples of the user speaking the hotword.

14.

发明申请
RECOGNIZING SPEECH IN THE PRESENCE OF ADDITIONAL AUDIO 有权

公开(公告)号：US20210272562A1

公开(公告)日：2021-09-02

申请号：US17303139

申请日：2021-05-21

Applicant: Google LLC

Inventor： Diego Melendo Casado , Ignacio Lopez Moreno , Javier Gonzalez-Dominguez

IPC: G10L15/20 , G06F3/16 , H03G3/30 , G10L15/22 , G10L17/06 , G10L21/034 , G10L25/84

Abstract: The technology described in this document can be embodied in a computer-implemented method that includes receiving, at a processing system, a first signal including an output of a speaker device and an additional audio signal. The method also includes determining, by the processing system, based at least in part on a model trained to identify the output of the speaker device, that the additional audio signal corresponds to an utterance of a user. The method further includes initiating a reduction in an audio output level of the speaker device based on determining that the additional audio signal corresponds to the utterance of the user.

15.

发明申请
Synthesis of Speech from Text in a Voice of a Target Speaker Using Neural Networks 有权

公开(公告)号：US20210217404A1

公开(公告)日：2021-07-15

申请号：US17055951

申请日：2019-05-17

Applicant: Google LLC

Inventor： Ye Jia , Zhifeng Chen , Yonghui Wu , Jonathan Shen , Ruoming Pang , Ron J. Weiss , Ignacio Lopez Moreno , Fei Ren , Yu Zhang , Quan Wang , Patrick Nguyen

IPC: G10L13/04 , G10L19/00 , G10L17/04

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech synthesis. The methods, systems, and apparatus include actions of obtaining an audio representation of speech of a target speaker, obtaining input text for which speech is to be synthesized in a voice of the target speaker, generating a speaker vector by providing the audio representation to a speaker encoder engine that is trained to distinguish speakers from one another, generating an audio representation of the input text spoken in the voice of the target speaker by providing the input text and the speaker vector to a spectrogram generation engine that is trained using voices of reference speakers to generate audio representations, and providing the audio representation of the input text spoken in the voice of the target speaker for output.

16.

发明申请
AUTOMATICALLY DETERMINING LANGUAGE FOR SPEECH RECOGNITION OF SPOKEN UTTERANCE RECEIVED VIA AN AUTOMATED ASSISTANT INTERFACE 有权

公开(公告)号：US20210097981A1

公开(公告)日：2021-04-01

申请号：US17120906

申请日：2020-12-14

Applicant: Google LLC

Inventor： Pu-sen Chao , Diego Melendo Casado , Ignacio Lopez Moreno

IPC: G10L15/14 , G10L15/02 , G10L15/18 , G06F3/16 , G10L15/00 , G10L15/183 , G10L15/22 , G10L15/30

Abstract: Implementations relate to determining a language for speech recognition of a spoken utterance, received via an automated assistant interface, for interacting with an automated assistant. Implementations can enable multilingual interaction with the automated assistant, without necessitating a user explicitly designate a language to be utilized for each interaction. Selection of a speech recognition model for a particular language can based on one or more interaction characteristics exhibited during a dialog session between a user and an automated assistant. Such interaction characteristics can include anticipated user input types, anticipated user input durations, a duration for monitoring for a user response, and/or an actual duration of a provided user response.

17.

发明申请
ADAPTIVE INTERFACE IN A VOICE-BASED NETWORKED SYSTEM 审中-公开

公开(公告)号：US20190318729A1

公开(公告)日：2019-10-17

申请号：US15973461

申请日：2018-05-07

Applicant: Google LLC

Inventor： Pu-sen Chao , Diego Melendo Casado , Ignacio Lopez Moreno

IPC: G10L15/18 , G10L15/08

Abstract: Determining a language for speech recognition of a spoken utterance received via an automated assistant interface for interacting with an automated assistant. The system can enable multilingual interaction with the automated assistant, without necessitating a user explicitly designate a language to be utilized for each interaction. The system can determine a user profile that corresponds to audio data that captures a spoken utterance, and utilize language(s), and optionally corresponding probabilities, assigned to the user profile in determining a language for speech recognition of the spoken utterance. The system can perform speech recognition in each of multiple languages assigned to the user profile, and utilize criteria to select only one of the speech recognitions as appropriate for generating and providing content that is responsive to the spoken utterance.

18.

发明授权
Improving speaker verification across locations, languages, and/or dialects 有权

公开(公告)号：US10403291B2

公开(公告)日：2019-09-03

申请号：US15995480

申请日：2018-06-01

Applicant: Google LLC

Inventor： Ignacio Lopez Moreno , Li Wan , Quan Wang

IPC: G10L17/24 , G10L17/02 , G10L17/08 , G10L17/14 , G10L17/18 , G10L17/22

Abstract: Methods, systems, apparatus, including computer programs encoded on computer storage medium, to facilitate language independent-speaker verification. In one aspect, a method includes actions of receiving, by a user device, audio data representing an utterance of a user. Other actions may include providing, to a neural network stored on the user device, input data derived from the audio data and a language identifier. The neural network may be trained using speech data representing speech in different languages or dialects. The method may include additional actions of generating, based on output of the neural network, a speaker representation and determining, based on the speaker representation and a second representation, that the utterance is an utterance of the user. The method may provide the user with access to the user device based on determining that the utterance is an utterance of the user.

19.

发明申请
MULTI-USER AUTHENTICATION ON A DEVICE 审中-公开

公开(公告)号：US20180308491A1

公开(公告)日：2018-10-25

申请号：US15956350

申请日：2018-04-18

Applicant: Google LLC

Inventor： Meltem Oktem , Taral Pradeep Joglekar , Fnu Heryandi , Pu-sen Chao , Ignacio Lopez Moreno , Salil Rajadhyaksha , Alexander H. Gruenstein , Diego Melendo Casado

IPC: G10L17/06 , G10L15/07 , G06F21/32 , G10L15/08 , G06F17/30 , G06K9/00

Abstract: In some implementations, authentication tokens corresponding to known users of a device are stored on the device. An utterance from a speaker is received. The utterance is classified as spoken by a particular known user of the known users. A query that includes a representation of the utterance and an indication of the particular known user as the speaker is provided using the authentication token of the particular known user.

20.

发明申请
RECOGNIZING SPEECH IN THE PRESENCE OF ADDITIONAL AUDIO 审中-公开

公开(公告)号：US20180211653A1

公开(公告)日：2018-07-26

申请号：US15887034

申请日：2018-02-02

Applicant: Google LLC

Inventor： Diego Melendo Casado , Ignacio Lopez Moreno , Javier Gonzalez-Dominguez

IPC: G10L15/20 , G10L25/84 , G10L21/034 , G10L15/26

CPC classification number: G10L15/222 , G06F3/165 , G06F3/167 , G10L15/20 , G10L15/265 , G10L17/00 , G10L17/06 , G10L21/034 , G10L25/84 , H03G3/3005

Abstract: The technology described in this document can be embodied in a computer-implemented method that includes receiving, at a processing system, a first signal including an output of a speaker device and an additional audio signal. The method also includes determining, by the processing system, based at least in part on a model trained to identify the output of the speaker device, that the additional audio signal corresponds to an utterance of a user. The method further includes initiating a reduction in an audio output level of the speaker device based on determining that the additional audio signal corresponds to the utterance of the user.

Search Results

Country/Region

Patent validity

Application date

Publication (announcement) day

applicant

The country/region where the applicant is located

Inventor

IPC

IPC Department

IPC class

IPC subclass

IPC group

IPC team

Appearance classification