摘要:
A method, information system, and computer-readable medium is provided for segmenting a plurality of data, such as multimedia data, and in particular an image document stream. Segment boundary points may be used for retrieving and/or browsing the plurality of data. Similarly, segment boundary points may be used to summarize the plurality of data. Examples of image document streams include video, PowerPoint slides, and NoteLook pages. A genetic method having a fitness or evaluation function using information retrieval concepts, such as importance and precedence, is used to obtain segment boundary points. The genetic method is able to evaluate a large amount of data in a cost effective manner. The genetic method is also able to run incrementally on streaming video and adapt to usage patterns by considering frequently accessed images.
摘要:
Techniques are provided for determining relevant information from a document based on document structure. A document is selected and structural elements within the document having a dominance relationship are determined. A first location within the document is selected. The structural element surrounding the first location is determined and the surrounding and non-surrounding structural elements are characterized. Additional documents are associated with the first location in the surrounding structural element based on the surrounding structural element characterization and the non-surrounding structural element characterization. Techniques for dynamically determining annotations for images based on document structure are also provided.
摘要:
The present invention relates to a method to make effective use of non rectangular display space for displaying a collage. In an embodiment of the invention, a heterogeneous set of images can be arranged to display the region of interest of the images to avoid overlapping regions of interest. The background gaps between the regions of interest can be filled by extending the regions of interest using a Voronoi technique. This produces a stained glass effect for the collage. In an embodiment of the present invention, the technique can be applied to irregular shapes including circular shapes with a hole in the middle. In an embodiment of the present invention, the technique can be used to print labels for disks.
摘要:
The present invention relates to a method to make effective use of non rectangular display space for displaying a collage. In an embodiment of the invention, a heterogeneous set of images can be arranged to display the region of interest of the images to avoid overlapping regions of interest. The background gaps between the regions of interest can be filled by extending the regions of interest using a Voronoi technique. This produces a stained glass effect for the collage. In an embodiment of the present invention, the technique can be applied to irregular shapes including circular shapes with a hole in the middle. In an embodiment of the present invention, the technique can be used to print labels for disks.
摘要:
Video recordings of meetings and scanned paper documents are natural digital documents that come out of a meeting. These can be placed on the Internet for easy access, with links generated between them by matching scanned documents to a segment of the video referencing the scanned document. Furthermore, annotations made on the paper documents during the meeting can be extracted and used as indexes to the video. An orthonormal transform, such as a Digital Cosine Transform (DCT) is used to compare scanned documents to video frames.
摘要:
In one embodiment, the present invention extracts video regions of interest from one or more videos and generates a highly condensed visual summary of the videos. The video regions of interest are extracted based on to energy, movement, face or other object detection methods, associated data or external input, or some other feature of the video. In another embodiment, the present invention extracts regions of interest from images and generates highly condensed visual summaries of the images. The highly condensed visual summary is generated by laying out germs on a canvas and then filling the spaces between the germs. The result is a visual summary that resembles a stained glass window having cells of varying shape. The germs may be laid out by temporal order, color histogram, similarity, according to a desired pattern, size, or some other manner. The people, objects and other visual content in the germs appear larger and become easier to see. The visual summary of the present invention utilizes important regions within the key frames, leading to more condensed summaries that are well suitable for small screens.
摘要:
In one embodiment, the present invention extracts video regions of interest from one or more videos and generates a highly condensed visual summary of the videos. The video regions of interest are extracted based on to energy, movement, face or other object detection methods, associated data or external input, or some other feature of the video. In another embodiment, the present invention extracts regions of interest from images and generates highly condensed visual summaries of the images. The highly condensed visual summary is generated by laying out germs on a canvas and then filling the spaces between the germs. The result is a visual summary that resembles a stained glass window having cells of varying shape. The germs may be laid out by temporal order, color histogram, similarity, according to a desired pattern, size, or some other manner. The people, objects and other visual content in the germs appear larger and become easier to see. The visual summary of the present invention utilizes important regions within the key frames, leading to more condensed summaries that are well suitable for small screens.
摘要:
A 3D graphical user interface includes a two-dimensional ground-plane layout representing the relationship between one or more leaf elements of a tree data structure. The interface further includes at least one building-like structure, each of the at least one building-like structure corresponding to a respective one of the one or more leaf elements. Each of the at least one building-like structure provides a summary of media associated with the respective one of the more leaf elements corresponding to at least one building-like structure.
摘要:
An algorithm for finding regions of interest (ROI) in synthetic images based on an information driven approach in which sub-blocks of a set of synthetic image are analyzed for information content or compressibility based on textural and color features. A DCT may be used to analyze the textural features of a set of images and a color histogram may be used to analyze the color features of the set of images. Sub-blocks of low compressibility are grouped into ROIs using a type of morphological technique. Unlike other algorithms that are geared for highly specific types of ROI (e.g. OCR text detection), the method of the present invention is generally applicable to arbitrary synthetic images. The present invention can be used with several other image applications, including Stained-Glass collages presentations.
摘要:
An algorithm for finding regions of interest (ROI) in images and photos based on an information driven approach in which sub-blocks of an image are analyzed for information content or compressibility based on the discrete cosine transform. The sub-blocks of low compressibility are grouped into ROIs using a morphological technique. Unlike other algorithms that are geared for highly specific types of ROI (e.g. face detection), the method of the present invention is generally applicable to arbitrary images and photos. A center-weighted variation of the algorithm can produce better results for certain photo applications. The algorithm can be used with several other image applications, including Stained-Glass collages and Pan-and-Scan presentations.