- 专利标题: AUTOMATIC SMOOTHED CAPTIONING OF NON-SPEECH SOUNDS FROM AUDIO
-
申请号: US15245152申请日: 2016-08-23
-
公开(公告)号: US20170278525A1公开(公告)日: 2017-09-28
- 发明人: Fangzhou Wang , Sourish Chaudhuri , Daniel Ellis , Nathan Reale
- 申请人: Google Inc.
- 主分类号: G10L21/10
- IPC分类号: G10L21/10 ; G10L15/20 ; G06F17/24 ; G10L25/84
摘要:
A content server accessing an audio stream, and inputs portions of the audio stream into one or more non-speech classifiers for classification, the non-speech classifiers generating, for portions of the audio stream, a set of raw scores representing likelihoods that the respective portion of the audio stream includes an occurrence of a particular class of non-speech sounds associated with each of the non-speech classifiers. The content server generates binary scores for the sets of raw scores, the binary scores generated based on a smoothing of a respective set of raw scores. The content server applies a set of non-speech captions to portions of the audio stream in time, each of the sets of non-speech captions based on a different one of the set binary scores of the corresponding portion of the audio stream.
公开/授权文献
信息查询