- 专利标题: Voice shortcut detection with speaker verification
-
申请号: US18103324申请日: 2023-01-30
-
公开(公告)号: US12033641B2公开(公告)日: 2024-07-09
- 发明人: Rajeev Rikhye , Quan Wang , Yanzhang He , Qiao Liang , Ian C. McGraw
- 申请人: Google LLC
- 申请人地址: US CA Mountain View
- 专利权人: GOOGLE LLC
- 当前专利权人: GOOGLE LLC
- 当前专利权人地址: US CA Mountain View
- 代理商 Gray Ice Higdon
- 主分类号: G10L17/24
- IPC分类号: G10L17/24 ; G10L15/26 ; G10L17/06 ; G10L21/028
摘要:
Techniques disclosed herein are directed towards streaming keyphrase detection which can be customized to detect one or more particular keyphrases, without requiring retraining of any model(s) for those particular keyphrase(s). Many implementations include processing audio data using a speaker separation model to generate separated audio data which isolates an utterance spoken by a human speaker from one or more additional sounds not spoken by the human speaker, and processing the separated audio data using a text independent speaker identification model to determine whether a verified and/or registered user spoke a spoken utterance captured in the audio data. Various implementations include processing the audio data and/or the separated audio data using an automatic speech recognition model to generate a text representation of the utterance. Additionally or alternatively, the text representation of the utterance can be processed to determine whether at least a portion of the text representation of the utterance captures a particular keyphrase. When the system determines the registered and/or verified user spoke the utterance and the system determines the text representation of the utterance captures the particular keyphrase, the system can cause a computing device to perform one or more actions corresponding to the particular keyphrase.
公开/授权文献
- US20230169984A1 VOICE SHORTCUT DETECTION WITH SPEAKER VERIFICATION 公开/授权日:2023-06-01
信息查询