-
公开(公告)号:US20240135973A1
公开(公告)日:2024-04-25
申请号:US17967364
申请日:2022-10-17
Applicant: Adobe Inc.
Inventor: Xue BAI , Justin Jonathan SALAMON , Aseem Omprakash AGARWALA , Hijung SHIN , Haoran CAI , Joel Richard BRANDT , Lubomira Assenova DONTCHEVA , Cristin Ailidh Fraser
IPC: G11B27/036 , G06F40/166 , G10L15/26 , G10L25/57 , G11B27/34
CPC classification number: G11B27/036 , G06F40/166 , G10L15/26 , G10L25/57 , G11B27/34 , G06F3/0482
Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for identifying candidate boundaries for video segments, video segment selection using those boundaries, and text-based video editing of video segments selected via transcript interactions. In an example implementation, boundaries of detected sentences and words are extracted from a transcript, the boundaries are retimed into an adjacent speech gap to a location where voice or audio activity is a minimum, and the resulting boundaries are stored as candidate boundaries for video segments. As such, a transcript interface presents the transcript, interprets input selecting transcript text as an instruction to select a video segment with corresponding boundaries selected from the candidate boundaries, and interprets commands that are traditionally thought of as text-based operations (e.g., cut, copy, paste) as an instruction to perform a corresponding video editing operation using the selected video segment.
-
公开(公告)号:US20240257798A1
公开(公告)日:2024-08-01
申请号:US18104434
申请日:2023-02-01
Applicant: ADOBE INC.
Inventor: Oriol NIETO-CABALLERO , Zeyu JIN , Justin Jonathan SALAMON , Franck DERNONCOURT
CPC classification number: G10L15/005 , G10L25/30
Abstract: Some aspects of the technology described herein employ a neural network with an efficient and lightweight architecture to perform spoken language recognition. Given an audio signal comprising speech, features are generated from the audio signal, for instance, by converting the audio signal to a normalized spectrogram. The features are input to the neural network, which has one or more convolutional layers and an output activation layer. Each neuron of the output activation layer corresponds to a language from a set of language and generates an activation value. Based on the activations values, an indication of zero or more languages from the set of languages is provided for the audio signal.
-
公开(公告)号:US20240127820A1
公开(公告)日:2024-04-18
申请号:US17967502
申请日:2022-10-17
Applicant: Adobe Inc.
Inventor: Justin Jonathan SALAMON , Fabian David CABA HEILBRON , Xue BAI , Aseem Omprakash AGARWALA , Hijung SHIN , Lubomira Assenova DONTCHEVA
IPC: G10L15/26 , G11B27/031
CPC classification number: G10L15/26 , G11B27/031
Abstract: Embodiments of the present invention provide systems, methods, and computer storage media for music-aware speaker diarization. In an example embodiment, one or more audio classifiers detect speech and music independently of each other, which facilitates detecting regions in an audio track that contain music but do not contain speech. These music-only regions are compared to the transcript, and any transcription and speakers that overlap in time with the music-only regions are removed from the transcript. In some embodiments, rather than having the transcript display the text from this detected music, a visual representation of the audio waveform is included in the corresponding regions of the transcript.
-
-