摘要:
A method and apparatus for adding new learning tasks to an incremental supervised learner provide a flexible incremental representation of all training examples encountered, thereby permitting state representations for new learning tasks to take advantage of incremental training already completed by encoding all past training examples as negative examples for a hypothetical learning task. The state representation of the hypothetical learning task is copied as the initial state representation for a new learning task to be initiated, is initialized with negative training examples of all previously presented training examples, thereby permitting the learning task to incorporate the previous examples efficiently.
摘要:
A meta-search engine apparatus and method for searching distributed networks using a plurality of search devices. The meta-search engine apparatus sends search queries to a plurality of search engines and compiles the results obtained from each of these search engines into a single ranked list. The results obtained from each of the search engines includes a listing of the titles of found sources of the search terms, or related search terms, and a summary of the source. The compilation and ranking is based primarily on the occurrence of search terms, or related search terms, in the titles and summaries but may also be based on, for example, relative weights given to each search engine, the number of search engines returning the same source as a result of a search, weighting of sections of the results obtained from the search engines, and the like.
摘要:
Methods of document expansion for a speech retrieval document by a recognizer. A database of vectors of automatic transcriptions of documents is accessed and the vectors are truncated by removing all terms that are not recognizable by the recognizer to create truncated vectors. Terms in the vectors are then weighted to associate the truncated vectors with the untruncated vectors. Terms not recognized by the recognizer are then added back to the weighted, truncated vectors. The retrieval effectiveness may then be measured.
摘要:
Methods of document expansion for a speech retrieval document by a recognizer. A database of vectors of automatic transcriptions of documents is accessed and the vectors are truncated by removing all terms that are not recognizable by the recognizer to create truncated vectors. Terms in the vectors are then weighted to associate the truncated vectors with the untruncated vectors. Terms not recognized by the recognizer are then added back to the weighted, truncated vectors. The retrieval effectiveness may then be measured.
摘要:
A method stores, indexes, searches and retrieves data information in a large data storage and retrieval system. Large amounts of data information, subject to searching and retrieval, are broken down and stored in sub-collections. Each sub-collection separately performs indexing of only the data information contained within that sub-collection and forms an inverted index. Statistical information derived from the inverted index of each sub-collection is collected by a global collection custodian and compiled into a global index. The global index is then passed to each sub-collection and is used by each during searching and retrieving of data information. Search results from each sub-collection are passed to the global collection custodian and organized there before being passed to a system user.
摘要:
Apparatus for adding new learning tasks to an incremental supervised learner provides a flexible incremental representation of all encountered training examples, thereby permitting state representations for new learning tasks to take advantage of incremental training already completed by encoding all past training examples as negative examples for a hypothetical learning task. The state representation of the hypothetical learning task is copied as the initial state representation for a new learning task to be initiated, and is initialized with negative training examples of all previously presented training examples, thereby permitting the learning task to efficiently incorporate the previous examples.
摘要:
A method stores, indexes, searches and retrieves data information in a large data storage and retrieval system. Large amounts of data information, subject to searching and retrieval, are broken down and stored in sub-collections. Each sub-collection separately performs indexing of only the data information contained within that sub-collection and forms an inverted index. Statistical information derived from the inverted index of each sub-collection is collected by a global collection custodian and compiled into a global index. The global index is then passed to each sub-collection and is used by each during searching and retrieving of data information. Search results from each sub-collection are passed to the global collection custodian and organized there before being passed to a system user.
摘要:
A method stores, indexes, searches and retrieves data information in a large data storage and retrieval system. Large amounts of data information, subject to searching and retrieval, are broken down and stored in sub-collections. Each sub-collection separately performs indexing of only the data information contained within that sub-collection and forms an inverted index. Statistical information derived from the inverted index of each sub-collection is collected by a global collection custodian and compiled into a global index. The global index is then passed to each sub-collection and is used by each during searching and retrieving of data information. Search results from each sub-collection are passed to the global collection custodian and organized there before being passed to a system user.
摘要:
Methods of document expansion for a speech retrieval document by a recognizer. A database of vectors of automatic transcriptions of documents is accessed and the vectors are truncated by removing all terms that are not recognizable by the recognizer to create truncated vectors. Terms in the vectors are then weighted to associate the truncated vectors with the untruncated vectors. Terms not recognized by the recognizer are then added back to the weighted, truncated vectors. The retrieval effectiveness may then be measured.
摘要:
A method and apparatus for communicating accumulated state information between internal and external tasks in a supervised learning system. A supervised learning system encodes state information for a hypothetical learning task on initialization. This hypothetical learning task state information indicates that no training instances have been received. During the supervised learning, training instances are presented to the supervised learner. The training instances are encoded with feature vector and target value information. For each task name paired with a non-default target value, the learner initializes a new learning task by copying the hypothetical learning task state representation for use as the state representation for the new learning task. Predictors are then produced for all learning tasks, except the hypothetical learning task. The new training instance is used to update all learning tasks as specified in the target vector. The new training instance is then used.to update the hypothetical learning task state representation as a negative example. Further training instances are handled similarly, new learning tasks are started based on the examination of the sparse target vector for task name, target value pairs which match received training instance target values and for which tasks have not yet been started.