摘要:
Music videos are automatically produced from source audio and video signals. The music video contains edited portions of the video signal synchronized with the audio signal. An embodiment detects transition points in the audio signal and the video signal. The transition points are used to align in time the video and audio signals. The video signal is edited according to its alignment with the audio signal. The resulting edited video signal is merged with the audio signal to form a music video.
摘要:
An algorithm for finding regions of interest (ROI) in synthetic images based on an information driven approach in which sub-blocks of a set of synthetic image are analyzed for information content or compressibility based on textural and color features. A DCT may be used to analyze the textural features of a set of images and a color histogram may be used to analyze the color features of the set of images. Sub-blocks of low compressibility are grouped into ROIs using a type of morphological technique. Unlike other algorithms that are geared for highly specific types of ROI (e.g. OCR text detection), the method of the present invention is generally applicable to arbitrary synthetic images. The present invention can be used with several other image applications, including Stained-Glass collages presentations.
摘要:
A media browser, graphical user interface and method for browsing a media file wherein a user selects at least one feature in a media file and is provided with information regarding the existence of the selected feature in the media file. Based on the information, the user can identify and playback portions of interest in a media file. Features in a media file, such as a speaker's identity, applause, silence, motion, or video cuts, are preferably automatically time-wise evaluated in the media file using known methods. Metadata generated based on the time-wise feature evaluation are preferably mapped to confidence score values that represent a probability of a corresponding feature's existence in the media file. Confidence score information is preferably presented graphically to a user as part of a graphical user interface, and is used to interactively browse the media file.
摘要:
An adaptive, interactive visual workspace for viewing groups of files based on their relationships. Relationships of files are visualized using iterative refinement of categories through a direct-manipulation graph-based layout. The visual workspace starts with a fully connected graph linking thumbnail images of related files that is then partitioned into neighborhoods in response to a user creating file stacks corresponding to different categories. Normalized spring lengths improve the overall quality of the layout. Different modes for membership in neighborhoods avoid confusing motion of files and help a user to manually organize the workspace. Additionally, retrieved files can be added without having to significantly move the previous files. Different visualization techniques indicate which files are related to each other. Different zoom rates are used for file location, and surrogate sizes allow users to increase the separation between files while still increasing the surrogate sizes.
摘要:
The present invention relates to a method to make effective use of non rectangular display space for displaying a collage. In an embodiment of the invention, a heterogeneous set of images can be arranged to display the region of interest of the images to avoid overlapping regions of interest. The background gaps between the regions of interest can be filled by extending the regions of interest using a Voronoi technique. This produces a stained glass effect for the collage. In an embodiment of the present invention, the technique can be applied to irregular shapes including circular shapes with a hole in the middle. In an embodiment of the present invention, the technique can be used to print labels for disks.
摘要:
The present invention relates to a method to make effective use of non rectangular display space for displaying a collage. In an embodiment of the invention, a heterogeneous set of images can be arranged to display the region of interest of the images to avoid overlapping regions of interest. The background gaps between the regions of interest can be filled by extending the regions of interest using a Voronoi technique. This produces a stained glass effect for the collage. In an embodiment of the present invention, the technique can be applied to irregular shapes including circular shapes with a hole in the middle. In an embodiment of the present invention, the technique can be used to print labels for disks.
摘要:
Video recordings of meetings and scanned paper documents are natural digital documents that come out of a meeting. These can be placed on the Internet for easy access, with links generated between them by matching scanned documents to a segment of the video referencing the scanned document. Furthermore, annotations made on the paper documents during the meeting can be extracted and used as indexes to the video. An orthonormal transform, such as a Digital Cosine Transform (DCT) is used to compare scanned documents to video frames.
摘要:
In one embodiment, the present invention extracts video regions of interest from one or more videos and generates a highly condensed visual summary of the videos. The video regions of interest are extracted based on to energy, movement, face or other object detection methods, associated data or external input, or some other feature of the video. In another embodiment, the present invention extracts regions of interest from images and generates highly condensed visual summaries of the images. The highly condensed visual summary is generated by laying out germs on a canvas and then filling the spaces between the germs. The result is a visual summary that resembles a stained glass window having cells of varying shape. The germs may be laid out by temporal order, color histogram, similarity, according to a desired pattern, size, or some other manner. The people, objects and other visual content in the germs appear larger and become easier to see. The visual summary of the present invention utilizes important regions within the key frames, leading to more condensed summaries that are well suitable for small screens.
摘要:
In one embodiment, the present invention extracts video regions of interest from one or more videos and generates a highly condensed visual summary of the videos. The video regions of interest are extracted based on to energy, movement, face or other object detection methods, associated data or external input, or some other feature of the video. In another embodiment, the present invention extracts regions of interest from images and generates highly condensed visual summaries of the images. The highly condensed visual summary is generated by laying out germs on a canvas and then filling the spaces between the germs. The result is a visual summary that resembles a stained glass window having cells of varying shape. The germs may be laid out by temporal order, color histogram, similarity, according to a desired pattern, size, or some other manner. The people, objects and other visual content in the germs appear larger and become easier to see. The visual summary of the present invention utilizes important regions within the key frames, leading to more condensed summaries that are well suitable for small screens.
摘要:
A 3D graphical user interface includes a two-dimensional ground-plane layout representing the relationship between one or more leaf elements of a tree data structure. The interface further includes at least one building-like structure, each of the at least one building-like structure corresponding to a respective one of the one or more leaf elements. Each of the at least one building-like structure provides a summary of media associated with the respective one of the more leaf elements corresponding to at least one building-like structure.