Speech enhancement machine learning model for estimation of reverberation in a multi-task learning framework

Invention Grant

US12014748B1 Speech enhancement machine learning model for estimation of reverberation in a multi-task learning framework 有权

Please log in to see more content

Patent Title: Speech enhancement machine learning model for estimation of reverberation in a multi-task learning framework
Application No.: US16988423

Application Date: 2020-08-07
Publication No.: US12014748B1

Publication Date: 2024-06-18
Inventor: Ritwik Giri , Mehmet Umut Isik , Neerad Dilip Phansalkar , Jean-Marc Valin , Karim Helwani , Arvindh Krishnaswamy
Applicant: Amazon Technologies, Inc.
Applicant Address: US WA Seattle
Assignee: Amazon Technologies, Inc.
Current Assignee: Amazon Technologies, Inc.
Current Assignee Address: US WA Seattle
Agency: NICHOLSON DE VOS WEBSTER & ELLIOTT LLP
Main IPC: G10L21/0208
IPC: G10L21/0208 ; G06N5/04 ; G06N20/00 ; G10L21/034

Speech enhancement machine learning model for estimation of reverberation in a multi-task learning framework

Abstract:

Techniques for training and using a machine learning model for estimation of reverberation in a multi-task learning framework are described. According to some embodiments, the multi-task learning framework improves the performance of the machine learning model by estimating the amount of reverberation present in an input audio recording as a secondary task to the primary task of generating a clean speech portion of the input audio recording. In one embodiment, a model architecture is selected that takes a noisy reverberant recording as an input and outputs an estimate of a clean (e.g., de-reverberated) signal, an estimate of noise (e.g., background noise), and an estimate of the reverb only portion, with the secondary task of estimating the reverb only portion acting as a regularizer that improves the machine learning model's performance in enhancing the reverberant (e.g., and noisy) input speech.

Information query

Espacenet

IPC分类:

G	物理
G10	乐器；声学
G10L	语音分析或合成；语音识别；语音或声音处理；语音或音频编码或解码
G10L21/00	为了改变语音或声音信号的质量或其可识度而处理语音或声音信号，以产生另一种可听的或非可听的信号，例如视觉信号或触觉信号（G10L19/00优先）
G10L21/02	.语音增强，例如降低噪声或消除回声（在直线传送系统中减轻回声效应入H04B3/20；免提电话中的回声抑制入H04M9/08）
G10L21/0208	..噪声过滤