Patent search ap:("Google LLC") AND inv:"Ignacio Lopez Moreno" Page 10

91.

发明授权
Neural networks for speaker verification 有权

公开(公告)号：US09978374B2

公开(公告)日：2018-05-22

申请号：US14846187

申请日：2015-09-04

Applicant: Google LLC

Inventor： Georg Heigold , Samy Bengio , Ignacio Lopez Moreno

IPC: G10L17/18 , G10L17/04 , G10L17/02

CPC classification number: G10L17/18 , G10L17/02 , G10L17/04

Abstract: This document generally describes systems, methods, devices, and other techniques related to speaker verification, including (i) training a neural network for a speaker verification model, (ii) enrolling users at a client device, and (iii) verifying identities of users based on characteristics of the users' voices. Some implementations include a computer-implemented method. The method can include receiving, at a computing device, data that characterizes an utterance of a user of the computing device. A speaker representation can be generated, at the computing device, for the utterance using a neural network on the computing device. The neural network can be trained based on a plurality of training samples that each: (i) include data that characterizes a first utterance and data that characterizes one or more second utterances, and (ii) are labeled as a matching speakers sample or a non-matching speakers sample.

92.

发明申请
TEXT INDEPENDENT SPEAKER RECOGNITION 有权

公开(公告)号：US20250131916A1

公开(公告)日：2025-04-24

申请号：US18965481

申请日：2024-12-02

Applicant: GOOGLE LLC

Inventor： Pu-sen Chao , Diego Melendo Casado , Ignacio Lopez Moreno , Quan Wang

IPC: G10L15/06 , G10L15/07 , G10L15/22 , G10L15/32 , G10L17/24

Abstract: Text independent speaker recognition models can be utilized by an automated assistant to verify a particular user spoke a spoken utterance and/or to identify the user who spoke a spoken utterance. Implementations can include automatically updating a speaker embedding for a particular user based on previous utterances by the particular user. Additionally or alternatively, implementations can include verifying a particular user spoke a spoken utterance using output generated by both a text independent speaker recognition model as well as a text dependent speaker recognition model. Furthermore, implementations can additionally or alternatively include prefetching content for several users associated with a spoken utterance prior to determining which user spoke the spoken utterance.

93.

发明授权
Automatically determining language for speech recognition of spoken utterance received via an automated assistant interface 有权

公开(公告)号：US12249319B2

公开(公告)日：2025-03-11

申请号：US18389033

申请日：2023-11-13

Applicant: GOOGLE LLC

Inventor： Pu-Sen Chao , Diego Melendo Casado , Ignacio Lopez Moreno

IPC: G10L15/14 , G06F3/16 , G10L15/00 , G10L15/02 , G10L15/08 , G10L15/18 , G10L15/183 , G10L15/22 , G10L15/30

Abstract: Implementations relate to determining a language for speech recognition of a spoken utterance, received via an automated assistant interface, for interacting with an automated assistant. Implementations can enable multilingual interaction with the automated assistant, without necessitating a user explicitly designate a language to be utilized for each interaction. Selection of a speech recognition model for a particular language can based on one or more interaction characteristics exhibited during a dialog session between a user and an automated assistant. Such interaction characteristics can include anticipated user input types, anticipated user input durations, a duration for monitoring for a user response, and/or an actual duration of a provided user response.

94.

发明公开
On-Device Multilingual Speech Recognition 审中-公开

公开(公告)号：US20240331700A1

公开(公告)日：2024-10-03

申请号：US18191711

申请日：2023-03-28

Applicant: Google LLC

Inventor： Yang Yu , Quan Wang , Ignacio Lopez Moreno

IPC: G10L15/26 , G10L15/32

CPC classification number: G10L15/26 , G10L15/32

Abstract: A method includes receiving a sequence of input audio frames and processing each corresponding input audio frame to determine a language ID event that indicates a predicted language. The method also includes obtaining speech recognition events each including a respective speech recognition result determined by a first language pack. Based on determining that the utterance includes a language switch from the first language to a second language, the method also includes loading a second language pack onto the client device and rewinding the input audio data buffered by an audio buffer to a time of the corresponding input audio frame associated with the language ID event that first indicated the second language as the predicted language. The method also includes emitting a first transcription and processing, using the second language pack loaded onto the client device, the rewound buffered audio data to generate a second transcription.

95.

发明公开
SPEAKER AWARENESS USING SPEAKER DEPENDENT SPEECH MODEL(S) 审中-公开

公开(公告)号：US20240203400A1

公开(公告)日：2024-06-20

申请号：US18394632

申请日：2023-12-22

Applicant: GOOGLE LLC

Inventor： Ignacio Lopez Moreno , Quan Wang , Jason Pelecanos , Li Wan , Alexander Gruenstein , Hakan Erdogan

IPC: G10L15/06 , G10L15/07 , G10L15/08 , G10L15/20 , G10L17/04 , G10L17/20 , G10L21/0208

CPC classification number: G10L15/063 , G10L15/07 , G10L15/20 , G10L17/04 , G10L17/20 , G10L21/0208 , G10L2015/088

Abstract: Implementations relate to an automated assistant that can bypass invocation phrase detection when an estimation of device-to-device distance satisfies a distance threshold. The estimation of distance can be performed for a set of devices, such as a computerized watch and a cellular phone, and/or any other combination of devices. The devices can communicate ultrasonic signals between each other, and the estimated distance can be determined based on when the ultrasonic signals are sent and/or received by each respective device. When an estimated distance satisfies the distance threshold, the automated assistant can operate as if the user is holding onto their cellular phone while wearing their computerized watch. This scenario can indicate that the user may be intending to hold their device to interact with the automated assistant and, based on this indication, the automated assistant can temporarily bypass invocation phrase detection (e.g., invoke the automated assistant).

96.

发明公开
AUTOMATICALLY DETERMINING LANGUAGE FOR SPEECH RECOGNITION OF SPOKEN UTTERANCE RECEIVED VIA AN AUTOMATED ASSISTANT INTERFACE 审中-公开

公开(公告)号：US20240054997A1

公开(公告)日：2024-02-15

申请号：US18382886

申请日：2023-10-23

Applicant: GOOGLE LLC

Inventor： Pu-sen Chao , Diego Melendo Casado , Ignacio Lopez Moreno

IPC: G10L15/197 , G10L15/00 , G10L15/22 , G10L15/30 , G10L15/08 , G10L15/14 , G10L15/18 , G10L13/00

CPC classification number: G10L15/197 , G10L15/005 , G10L15/22 , G10L15/30 , G10L15/08 , G10L15/14 , G10L15/1822 , G10L13/00 , G10L2015/088 , G10L2015/223 , G10L2015/228

Abstract: Determining a language for speech recognition of a spoken utterance received via an automated assistant interface for interacting with an automated assistant. Implementations can enable multilingual interaction with the automated assistant, without necessitating a user explicitly designate a language to be utilized for each interaction. Implementations determine a user profile that corresponds to audio data that captures a spoken utterance, and utilize language(s), and optionally corresponding probabilities, assigned to the user profile in determining a language for speech recognition of the spoken utterance. Some implementations select only a subset of languages, assigned to the user profile, to utilize in speech recognition of a given spoken utterance of the user. Some implementations perform speech recognition in each of multiple languages assigned to the user profile, and utilize criteria to select only one of the speech recognitions as appropriate for generating and providing content that is responsive to the spoken utterance.

97.

发明授权
Speaker awareness using speaker dependent speech model(s) 有权

公开(公告)号：US11854533B2

公开(公告)日：2023-12-26

申请号：US17587424

申请日：2022-01-28

Applicant: GOOGLE LLC

Inventor： Ignacio Lopez Moreno , Quan Wang , Jason Pelecanos , Li Wan , Alexander Gruenstein , Hakan Erdogan

IPC: G10L15/16 , G10L15/06 , G10L15/07 , G10L15/20 , G10L17/04 , G10L17/20 , G10L21/0208 , G10L15/08

CPC classification number: G10L15/063 , G10L15/07 , G10L15/20 , G10L17/04 , G10L17/20 , G10L21/0208 , G10L2015/088

Abstract: Techniques disclosed herein enable training and/or utilizing speaker dependent (SD) speech models which are personalizable to any user of a client device. Various implementations include personalizing a SD speech model for a target user by processing, using the SD speech model, a speaker embedding corresponding to the target user along with an instance of audio data. The SD speech model can be personalized for an additional target user by processing, using the SD speech model, an additional speaker embedding, corresponding to the additional target user, along with another instance of audio data. Additional or alternative implementations include training the SD speech model based on a speaker independent speech model using teacher student learning.

98.

发明授权
Automatically determining language for speech recognition of spoken utterance received via an automated assistant interface 有权

公开(公告)号：US11817085B2

公开(公告)日：2023-11-14

申请号：US17120906

申请日：2020-12-14

Applicant: Google LLC

Inventor： Pu-Sen Chao , Diego Melendo Casado , Ignacio Lopez Moreno

IPC: G10L15/14 , G10L15/02 , G10L15/18 , G06F3/16 , G10L15/00 , G10L15/183 , G10L15/22 , G10L15/30 , G10L15/08

CPC classification number: G10L15/14 , G06F3/167 , G10L15/005 , G10L15/02 , G10L15/183 , G10L15/1822 , G10L15/22 , G10L15/30 , G10L2015/088 , G10L2015/223 , G10L2015/228

Abstract: Implementations relate to determining a language for speech recognition of a spoken utterance, received via an automated assistant interface, for interacting with an automated assistant. Implementations can enable multilingual interaction with the automated assistant, without necessitating a user explicitly designate a language to be utilized for each interaction. Selection of a speech recognition model for a particular language can based on one or more interaction characteristics exhibited during a dialog session between a user and an automated assistant. Such interaction characteristics can include anticipated user input types, anticipated user input durations, a duration for monitoring for a user response, and/or an actual duration of a provided user response.

99.

发明授权
Speaker diarization using speaker embedding(s) and trained generative model 有权

公开(公告)号：US11735176B2

公开(公告)日：2023-08-22

申请号：US17215129

申请日：2021-03-29

Applicant: Google LLC

Inventor： Ignacio Lopez Moreno , Luis Carlos Cobo Rus

IPC: G10L15/00 , G10L15/20 , G10L15/30 , G10L15/02 , G10L15/22 , G10L21/0208 , G10L15/06

CPC classification number: G10L15/20 , G10L15/02 , G10L15/22 , G10L15/30 , G10L15/063 , G10L21/0208

Abstract: Speaker diarization techniques that enable processing of audio data to generate one or more refined versions of the audio data, where each of the refined versions of the audio data isolates one or more utterances of a single respective human speaker. Various implementations generate a refined version of audio data that isolates utterance(s) of a single human speaker by generating a speaker embedding for the single human speaker, and processing the audio data using a trained generative model—and using the speaker embedding in determining activations for hidden layers of the trained generative model during the processing. Output is generated over the trained generative model based on the processing, and the output is the refined version of the audio data.

100.

发明授权
Multi-user authentication on a device 有权

公开(公告)号：US11721326B2

公开(公告)日：2023-08-08

申请号：US17584866

申请日：2022-01-26

Applicant: GOOGLE LLC

Inventor： Meltem Oktem , Taral Pradeep Joglekar , Fnu Heryandi , Pu-sen Chao , Ignacio Lopez Moreno , Salil Rajadhyaksha , Alexander H. Gruenstein , Diego Melendo Casado

IPC: G10L15/08 , G06F21/32 , G10L17/06 , G06F16/635 , G10L15/22 , G10L17/00 , G06V40/10 , G10L15/07 , G10L15/26

CPC classification number: G10L15/08 , G06F16/636 , G06F21/32 , G06V40/10 , G10L15/07 , G10L15/22 , G10L17/00 , G10L17/06 , G10L15/26 , G10L2015/088

Abstract: In some implementations, processor(s) can receive an utterance from a speaker, and determine whether the speaker is a known user of a user device or not a known user of the user device. The user device can be shared by a plurality of known users. Further, the processor(s) can determine whether the utterance corresponds to a personal request or non-personal request. Moreover, and in response to determining that the speaker is not a known user of the user device and in response to determining that the utterance corresponds to a non-personal request, the processor(s) can cause a response to the utterance to be provided for presentation to the speaker at the user device response to the utterance, or can cause an action to be performed by the user device responsive to the utterance.

Search Results

Country/Region

Patent validity

Application date

Publication (announcement) day

applicant

The country/region where the applicant is located

Inventor

IPC

IPC Department

IPC class

IPC subclass

IPC group

IPC team

Appearance classification