Patent search ap:"SoundHound Inc." Page 1

1.

发明授权
Deriving acoustic features and linguistic features from received speech audio 有权

公开(公告)号：US12175964B2

公开(公告)日：2024-12-24

申请号：US17325114

申请日：2021-05-19

Applicant: SoundHound, Inc.

Inventor： Kiran Garaga Lokeswarappa , Joel Gedalius , Bernard Mont-Reynaud , Jun Huang

IPC: G10L15/00 , G06F40/205 , G06F40/211 , G06F40/253 , G06N20/00 , G06Q30/0241 , G06Q30/0251 , G10L15/02 , G10L15/06 , G10L15/18 , G10L25/90 , H04L67/306 , G10L15/22 , G10L15/26 , G10L25/51

Abstract: A computer-implemented method is provided. The method including receiving speech audio of dictation associated with a user ID, deriving acoustic features from the speech audio, storing the derived acoustic features in a user profile associated with the user ID, receiving a request for acoustic features through an application programming interface (API), the request including the user ID, and sending the derived acoustic features through the API.

2.

发明授权
Content filtering in media playing devices 有权

公开(公告)号：US12126868B2

公开(公告)日：2024-10-22

申请号：US18348249

申请日：2023-07-06

Applicant: SoundHound, Inc.

Inventor： Thor S. Khov , Terry Kong

IPC: H04N21/454 , G06N3/045 , G06V20/40 , H04N21/44 , H04N21/466

CPC classification number: H04N21/4542 , G06N3/045 , G06V20/46 , H04N21/44008 , H04N21/4665

Abstract: Various approaches relate to user defined content filtering in media playing devices of undesirable content represented in stored and real-time content from content providers. For example, video, image, and/or audio data can be analyzed to identify and classify content included in the data using various classification models and object and text recognition approaches. Thereafter, the identification and classification can be used to control presentation and/or access to the content and/or portions of the content. For example, based on the classification, portions of the content can be modified (e.g., replaced, removed, degraded, etc.) using one or more techniques (e.g., media replacement, media removal, media degradation, etc.) and then presented.

3.

发明授权
Automatic learning of entities, words, pronunciations, and parts of speech 有权

公开(公告)号：US12080275B2

公开(公告)日：2024-09-03

申请号：US17146239

申请日：2021-01-11

Applicant: SoundHound, Inc.

Inventor： Anton V. Relin

IPC: G10L15/02 , G10L15/14 , G10L15/19

CPC classification number: G10L15/02 , G10L15/14 , G10L15/19 , G10L2015/025

Abstract: Systems for automatic speech recognition and/or natural language understanding automatically learn new words by finding subsequences of phonemes that, if they were a new word, would enable a successful tokenization of a phoneme sequence. Systems can learn alternate pronunciations of words by finding phoneme sequences with a small edit distance to existing pronunciations. Systems can learn the part of speech of words by finding part-of-speech variations that would enable parses by syntactic grammars. Systems can learn what types of entities a word describes by finding sentences that could be parsed by a semantic grammar but for the words not being on an entity list.

4.

发明授权
Method and system for conversation transcription with metadata 有权

公开(公告)号：US12020708B2

公开(公告)日：2024-06-25

申请号：US17450552

申请日：2021-10-11

Applicant: SoundHound, Inc.

Inventor： Kiersten L. Bradley , Ethan Coeytaux , Ziming Yin

IPC: G10L15/26 , G06F40/134 , G06F40/166 , G06F40/284 , G10L15/02 , G10L15/06 , G10L15/07

CPC classification number: G10L15/26 , G06F40/134 , G06F40/166 , G06F40/284 , G10L15/02 , G10L15/063 , G10L15/07 , G10L2015/0631

Abstract: Methods and systems for enabling an efficient review of meeting content via a metadata-enriched, speaker-attributed transcript are disclosed. By incorporating speaker diarization and other metadata, the system can provide a structured and effective way to review and/or edit the transcript. One type of metadata can be image or video data to represent the meeting content. Furthermore, the present subject matter utilizes a multimodal diarization model to identify and label different speakers. The system can synchronize various sources of data, e.g., audio channel data, voice feature vectors, acoustic beamforming, image identification, and extrinsic data, to implement speaker diarization.

5.

发明授权
Method for providing information, method for generating database, and program 有权

公开(公告)号：US11995143B2

公开(公告)日：2024-05-28

申请号：US17649052

申请日：2022-01-26

Applicant: SoundHound, Inc.

Inventor： Masaki Naito , Keisuke Tsuchida , Jun Yoneyama , Kaku Sawada

IPC: G06F16/95 , G06F16/33 , G06F16/955 , G06F40/40 , G10L15/26

CPC classification number: G06F16/9566 , G06F16/3344 , G06F40/40 , G10L15/26

Abstract: As audio (1) is input to an extension of a browser, the extension transmits the audio (1) to a language processing server. A speech recognition unit obtains a text (1) corresponding to the audio (1), and transmits the text (1) to a natural language understanding unit. In the natural language understanding unit, an information processing unit identifies a URL (1) corresponding to the text (1), and transmits the URL (1) to the browser. The extension passes the URL (1) to a browsing function. The browsing function uses the URL (1) to access a web server. The web server transmits a web page (1) corresponding to the URL (1) to the browser. The browsing function shows a screen corresponding to the web page (1) on a display.

6.

发明公开
DOMAIN SPECIFIC NEURAL SENTENCE GENERATOR FOR MULTI-DOMAIN VIRTUAL ASSISTANTS 审中-公开

公开(公告)号：US20240144921A1

公开(公告)日：2024-05-02

申请号：US18050182

申请日：2022-10-27

Applicant: SoundHound, Inc.

Inventor： Pranav SINGH , Yilun ZHANG , Eunjee NA , Olivia BETTAGLIO

IPC: G10L15/18 , G10L15/06 , G10L15/22

CPC classification number: G10L15/1815 , G10L15/063 , G10L15/1822 , G10L15/22 , G10L2015/0631 , G10L2015/223

Abstract: Automatically generating sentences that a user can say to invoke a set of defined actions performed by a virtual assistant are disclosed. A sentence is received and keywords are extracted from the sentence. Based on the keywords, additional sentences are generated. A classifier model is applied to the generated sentences to determine a sentence that satisfies a threshold. In the situation a sentence satisfies the threshold, an intent associated with the classifier model can be invoked. In the situation the sentences fail to satisfy the classifier model, the virtual assistant can attempt to interpret the received sentence according to the most likely intent by invoking a sentence generation model fine-tuned for a particular domain, generate additional sentences with a high probability of having the same intent and fulfill the specific action defined by the intent.

7.

发明公开
METHOD AND SYSTEM FOR PROACTIVE INTERACTION 审中-公开

公开(公告)号：US20240046923A1

公开(公告)日：2024-02-08

申请号：US18361791

申请日：2023-07-28

Applicant: SoundHound, Inc.

Inventor： Masaki NAITO

IPC: G10L15/19 , G06F16/245 , G10L15/22 , G10L15/30

CPC classification number: G10L15/19 , G06F16/245 , G10L15/22 , G10L15/30

Abstract: In an interaction system, a server can obtain a setting expression including a query and a condition for functioning as a virtual assistant, store the query and the condition in a memory, and deliver an inquiry expression including the query in response to occurrence of a situation specified by the condition. The setting expression can be by voice or natural language. Processes can be different for different users and can be based on domain. The inquiry expression includes a question asking the user for an affirmative response before performing the inquiry. Implementations can be adopted in or near a vehicle.

8.

发明公开
PRE-WAKEWORD SPEECH PROCESSING 审中-公开

公开(公告)号：US20230386458A1

公开(公告)日：2023-11-30

申请号：US17804544

申请日：2022-05-27

Applicant: SoundHound, Inc.

Inventor： Karl STAHL , Bernard MONT-REYNAUD

IPC: G10L15/22 , G10L15/08 , G10L25/93

CPC classification number: G10L15/22 , G10L15/08 , G10L25/93 , G10L2015/088

Abstract: Methods and systems for pre-wakeword speech processing are disclosed. Speech audio, comprising command speech spoken before a wakeword, may be stored in a buffer in oldest to newest order. Upon detection of the wakeword, reverse acoustic models and language models, such as reverse automatic speech recognition (R-ASR) can be applied to the buffered audio, in newest to oldest order, starting from before the wakeword. The speech is converted into a sequence of words. Natural language grammar models, such as natural language understanding (NLU), can be applied to match the sequence of words to a complete command, the complete command being associated with invoking a computer operation.

9.

发明授权
Training a device specific acoustic model 有权

公开(公告)号：US11830472B2

公开(公告)日：2023-11-28

申请号：US17573551

申请日：2022-01-11

Applicant: SOUNDHOUND, INC.

Inventor： Keyvan Mohajer , Mehul Patel

IPC: G10L15/22 , G06F3/16 , G10L15/18

CPC classification number: G10L15/22 , G06F3/167 , G10L15/18

Abstract: Developers can configure custom acoustic models by providing audio files with custom recordings. The custom acoustic model is trained by tuning a baseline model using the audio files. Audio files may contain custom noise to apply to clean speech for training. The custom acoustic model is provided as an alternative to a standard acoustic model. Device developers can select an acoustic model by a user interface. Speech recognition is performed on speech audio using one or more acoustic models. The result can be provided to developers through the user interface, and an error rate can be computed and also provided.

10.

发明公开
VIDEO CONFERENCE CAPTIONING 审中-公开

公开(公告)号：US20230245661A1

公开(公告)日：2023-08-03

申请号：US18298282

申请日：2023-04-10

Applicant: SoundHound, Inc.

Inventor： Ethan COEYTAUX

IPC: G10L15/26 , G10L15/02 , G10L19/005 , G10L15/19 , G10L15/14

CPC classification number: G10L15/26 , G10L15/02 , G10L19/005 , G10L15/19 , G10L15/14

Abstract: A video conferencing system, such as one implemented with a cloud server, receives audio streams from a plurality of endpoints. The system uses automatic speech recognition to transcribe speech in the audio streams. The system multiplexes the transcriptions into individual caption streams and sends them to the endpoints, but the caption stream to each endpoint omits the transcription of audio from the endpoint. Some systems allow muting of audio through an indication to the system. The system then omits sending the muted audio to other endpoints and also omits sending a transcription of the muted audio to other endpoints.

Search Results

Country/Region

Patent validity

Application date

Publication (announcement) day

applicant

The country/region where the applicant is located

Inventor

IPC

IPC Department

IPC class

IPC subclass

IPC group

IPC team

Appearance classification