摘要:
The present invention meets these needs by providing temporal coherency to recognition systems. One embodiment of the present invention comprises a manifold recognition module to use a sequence of images for recognition. A manifold training module receives a plurality of training image sequences (e.g. from a video camera), each training image sequence including an individual in a plurality of poses, and establishes relationships between the images of a training image sequence. A probabilistic identity module receives a sequence of recognition images including a target individual for recognition, and identifies the target individual based on the relationship of training images corresponding to the recognition images. An occlusion module masks occluded portions of an individual's face to prevent distorted identifications.
摘要:
Methods and systems are described for three-dimensional pose estimation. A training module determines a mapping function between a training image sequence and pose representations of a subject in the training image sequence. The training image sequence is represented by a set of appearance and motion patches. A set of filters are applied to the appearance and motion patches to extract features of the training images. Based on the extracted features, the training module learns a multidimensional mapping function that maps the motion and appearance patches to the pose representations of the subject. A testing module outputs a fast human pose estimation by applying the learned mapping function to a test image sequence.
摘要:
A system and method recognizes and tracks human motion from different motion classes. In a learning stage, a discriminative model is learned to project motion data from a high dimensional space to a low dimensional space while enforcing discriminance between motions of different motion classes in the low dimensional space. Additionally, low dimensional data may be clustered into motion segments and motion dynamics learned for each motion segment. In a tracking stage, a representation of human motion is received comprising at least one class of motion. The tracker recognizes and tracks the motion based on the learned discriminative model and the learned dynamics.
摘要:
Taking a set of unlabeled images of a collection of objects acquired under different imaging conditions, and decomposing the set into disjoint subsets corresponding to individual objects requires clustering. Appearance-based methods for clustering a set of images of 3-D objects acquired under varying illumination conditions can be based on the concept of illumination cones. A clustering problem is equivalent to finding convex polyhedral cones in the high-dimensional image space. To efficiently determine the conic structures hidden in the image data, the concept of conic affinity can be used which measures the likelihood of a pair of images belonging to the same underlying polyhedral cone. Other algorithms can be based on affinity measure based on image gradient comparisons operating directly on the image gradients by comparing the magnitudes and orientations of the image gradient.
摘要:
A system and a method are disclosed for an adaptive discriminative generative model with a probabilistic interpretation. As applied to visual tracking, the discriminative generative model separates the target object from the background more accurately and efficiently than conventional methods. A computationally efficient algorithm constantly updates the discriminative model over time. The discriminative generative model adapts to accommodate dynamic appearance variations of the target and background. Experiments show that the discriminative generative model effectively tracks target objects undergoing large pose and lighting changes.
摘要:
The advantage of the present invention is to appropriately detect the object. The object detection apparatus in the present invention has a plurality of cameras to determine the distance to the objects, a distance determination unit to determine the distance therein, a histogram generation unit to specify the frequency of the pixels against the distances to the pixels, an object distance determination unit that determines the most likely distance, a probability mapping unit that provides the probabilities of the pixels based on the difference of the distance, a kernel detection unit that determines a kernel region as a group of the pixels, a periphery detection unit that determines a peripheral region as a group of the pixels, selected from the pixels being close to the kernel region and an object specifying unit that specifies the object region where the object is present with a predetermined probability.
摘要:
A system and a method are disclosed for clustering images of objects seen from different viewpoints. That is, given an unlabelled set of images of n objects, an unsupervised algorithm groups the images into N disjoint subsets such that each subset only contains images of a single object. The clustering method makes use of a broad geometric framework that exploits the interplay between the geometry of appearance manifolds and the symmetry of the 2D affine group.
摘要:
A visual tracker tracks an object in a sequence of input images. A tracking module detects a location of the object based on a set of weighted blocks representing the object's shape. The tracking module then refines a segmentation of the object from the background image at the detected location. Based on the refined segmentation, the set of weighted blocks are updated. By adaptively encoding appearance and shape into the block configuration, the present invention is able to efficiently and accurately track an object even in the presence of rapid motion that causes large variations in appearance and shape of the object.
摘要:
Simultaneous localization and mapping (SLAM) utilizes multiple view feature descriptors to robustly determine location despite appearance changes that would stifle conventional systems. A SLAM algorithm generates a feature descriptor for a scene from different perspectives using kernel principal component analysis (KPCA). When the SLAM module subsequently receives a recognition image after a wide baseline change, it can refer to correspondences from the feature descriptor to continue map building and/or determine location. Appearance variations can result from, for example, a change in illumination, partial occlusion, a change in scale, a change in orientation, change in distance, warping, and the like. After an appearance variation, a structure-from-motion module uses feature descriptors to reorient itself and continue map building using an extended Kalman Filter. Through the use of a database of comprehensive feature descriptors, the SLAM module is also able to refine a position estimation despite appearance variations.
摘要:
A method and system efficiently and accurately detects humans in a test image and classifies their pose. In a training stage, a probabilistic model is derived in an unsupervised or semi-supervised manner such that at least some poses are not manually labeled. The model provides two sets of model parameters to describe the statistics of images containing humans and images of background scenes. In a testing stage, the probabilistic model is used to determine if a human is present in the image, and classify the human's pose based on the poses in the training images. A solution is efficiently provided to both human detection and pose classification by using the same probabilistic model to solve the problems.