Patent search ap:("SoundHound Page Inc.") AND inv:"Ethan COEYTAUX"

1.

发明公开
VIDEO CONFERENCE CAPTIONING 审中-公开

公开(公告)号：US20230245661A1

公开(公告)日：2023-08-03

申请号：US18298282

申请日：2023-04-10

Applicant: SoundHound, Inc.

Inventor： Ethan COEYTAUX

IPC: G10L15/26 , G10L15/02 , G10L19/005 , G10L15/19 , G10L15/14

CPC classification number: G10L15/26 , G10L15/02 , G10L19/005 , G10L15/19 , G10L15/14

Abstract: A video conferencing system, such as one implemented with a cloud server, receives audio streams from a plurality of endpoints. The system uses automatic speech recognition to transcribe speech in the audio streams. The system multiplexes the transcriptions into individual caption streams and sends them to the endpoints, but the caption stream to each endpoint omits the transcription of audio from the endpoint. Some systems allow muting of audio through an indication to the system. The system then omits sending the muted audio to other endpoints and also omits sending a transcription of the muted audio to other endpoints.

2.

发明申请
METHOD AND SYSTEM FOR CONVERSATION TRANSCRIPTION WITH METADATA 有权

公开(公告)号：US20220115020A1

公开(公告)日：2022-04-14

申请号：US17450552

申请日：2021-10-11

Applicant: SoundHound, Inc.

Inventor： Kiersten L. BRADLEY , Ethan COEYTAUX , Ziming YIN

IPC: G10L15/26 , G10L15/06 , G10L15/02

Abstract: Methods and systems for enabling an efficient review of meeting content via a metadata-enriched, speaker-attributed transcript are disclosed. By incorporating speaker diarization and other metadata, the system can provide a structured and effective way to review and/or edit the transcript. One type of metadata can be image or video data to represent the meeting content. Furthermore, the present subject matter utilizes a multimodal diarization model to identify and label different speakers. The system can synchronize various sources of data, e.g., audio channel data, voice feature vectors, acoustic beamforming, image identification, and extrinsic data, to implement speaker diarization.

3.

发明申请
METHOD AND SYSTEM FOR CONVERSATION TRANSCRIPTION WITH METADATA 有权

公开(公告)号：US20220115019A1

公开(公告)日：2022-04-14

申请号：US17450551

申请日：2021-10-11

Applicant: SoundHound, Inc.

Inventor： Kiersten L. BRADLEY , Ethan COEYTAUX , Ziming YIN

IPC: G10L15/26 , G10L15/07 , G06F40/166 , G06F40/284 , G06F40/134

Abstract: Methods and systems for enabling an efficient review of meeting content via a metadata-enriched, speaker-attributed and multiuser-editable transcript are disclosed. By incorporating speaker diarization and other metadata, the system can provide a structured and effective way to review and/or edit the transcript by one or more editors. One type of metadata can be image or video data to represent the meeting content. Furthermore, the present subject matter utilizes a multimodal diarization model to identify and label different speakers. The system can synchronize various sources of data, e.g., audio channel data, voice feature vectors, acoustic beamforming, image identification, and extrinsic data, to implement speaker diarization.

Patent Agency Ranking