Patent search ap:("Adobe Inc.") AND inv:"Justin Jonathan Salamon" Page 1

1.

发明公开
SYSTEMS AND METHODS FOR CROSS-MODAL RETRIEVAL BASED ON A SOUND MODALITY AND A NON-SOUND MODALITY 审中-公开

公开(公告)号：US20240362269A1

公开(公告)日：2024-10-31

申请号：US18308970

申请日：2023-04-28

Applicant: ADOBE INC.

Inventor： Ho-Hsiang Wu , Oriol Nieto , Justin Jonathan Salamon

IPC: G06F16/632 , G06F16/638 , G06F16/68

CPC classification number: G06F16/632 , G06F16/638 , G06F16/686

Abstract: Systems and methods for cross-modal retrieval are provided. According to one aspect, a method for cross-modal retrieval includes obtaining a query describing a sound using a query modality other than a sound modality; encoding the query to obtain a query embedding using a query encoder network for the query modality and a query projection network, wherein the query projection network includes a self-attention layer, and wherein the query embedding is in a joint embedding space for the query modality and the sound modality; and providing a response including an audio sample based on the query embedding, wherein the audio sample includes the sound.

2.

发明授权
Automatic recognition of visual and audio-visual cues 有权

公开(公告)号：US12125317B2

公开(公告)日：2024-10-22

申请号：US17539652

申请日：2021-12-01

Applicant: ADOBE INC.

Inventor： Jiyoung Lee , Justin Jonathan Salamon , Dingzeyu Li

IPC: G06V40/20 , G06N3/045 , G06N3/08 , G06V10/82 , G06V20/40

CPC classification number: G06V40/20 , G06N3/045 , G06N3/08 , G06V10/82 , G06V20/41 , G06V20/46 , G06V20/49

Abstract: A method for detecting a cue (e.g., a visual cue or a visual cue combined with an audible cue) occurring together in an input video includes: presenting a user interface to record an example video of a user performing an act including the cue; determining a part of the example video where the cue occurs; applying a feature of the part to a neural network to generate a positive embedding; dividing the input video into a plurality of chunks and applying a feature of each chunk to the neural network to generate a plurality of negative embeddings; applying a feature of a given one of the chunks to the neural network to output a query embedding; and determining whether the cue occurs in the input video from the query embedding, the positive embedding, and the negative embeddings.

3.

发明授权
Music-aware speaker diarization for transcripts and text-based video editing 有权

公开(公告)号：US12223962B2

公开(公告)日：2025-02-11

申请号：US17967502

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Justin Jonathan Salamon , Fabian David Caba Heilbron , Xue Bai , Aseem Omprakash Agarwala , Hijung Shin , Lubomira Assenova Dontcheva

IPC: G10L15/08 , G10L15/26 , G11B27/031

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for music-aware speaker diarization. In an example embodiment, one or more audio classifiers detect speech and music independently of each other, which facilitates detecting regions in an audio track that contain music but do not contain speech. These music-only regions are compared to the transcript, and any transcription and speakers that overlap in time with the music-only regions are removed from the transcript. In some embodiments, rather than having the transcript display the text from this detected music, a visual representation of the audio waveform is included in the corresponding regions of the transcript.

4.

发明授权
Video segment selection and editing using transcript interactions 有权

公开(公告)号：US12119028B2

公开(公告)日：2024-10-15

申请号：US17967364

申请日：2022-10-17

Applicant: Adobe Inc.

Inventor： Xue Bai , Justin Jonathan Salamon , Aseem Omprakash Agarwala , Hijung Shin , Haoran Cai , Joel Richard Brandt , Lubomira Assenova Dontcheva , Cristin Ailidh Fraser

IPC: G11B27/036 , G06F40/166 , G10L15/26 , G10L25/57 , G11B27/34 , G06F3/0482 , G06F3/04845 , G06F3/0485

CPC classification number: G11B27/036 , G06F40/166 , G10L15/26 , G10L25/57 , G11B27/34 , G06F3/0482 , G06F3/04845 , G06F3/0485

Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for identifying candidate boundaries for video segments, video segment selection using those boundaries, and text-based video editing of video segments selected via transcript interactions. In an example implementation, boundaries of detected sentences and words are extracted from a transcript, the boundaries are retimed into an adjacent speech gap to a location where voice or audio activity is a minimum, and the resulting boundaries are stored as candidate boundaries for video segments. As such, a transcript interface presents the transcript, interprets input selecting transcript text as an instruction to select a video segment with corresponding boundaries selected from the candidate boundaries, and interprets commands that are traditionally thought of as text-based operations (e.g., cut, copy, paste) as an instruction to perform a corresponding video editing operation using the selected video segment.

Search Results

Country/Region

Patent validity

Application date

Publication (announcement) day

applicant

The country/region where the applicant is located

Inventor

IPC

IPC Department

IPC class

IPC subclass

IPC group

IPC team

Appearance classification