发明申请
- 专利标题: SEARCH RESULTS RANKING USING EDITING DISTANCE AND DOCUMENT INFORMATION
- 专利标题(中): 搜索结果使用编辑距离和文档信息排名
-
申请号: US12101951申请日: 2008-04-11
-
公开(公告)号: US20090259651A1公开(公告)日: 2009-10-15
- 发明人: Vladimir Tankovich , Hang Li , Dmitriy Meyerzon , Jun Xu
- 申请人: Vladimir Tankovich , Hang Li , Dmitriy Meyerzon , Jun Xu
- 申请人地址: US WA Redmond
- 专利权人: MICROSOFT CORPORATION
- 当前专利权人: MICROSOFT CORPORATION
- 当前专利权人地址: US WA Redmond
- 主分类号: G06F17/30
- IPC分类号: G06F17/30
摘要:
Architecture for extracting document information from documents received as search results based on a query string, and computing an edit distance between the data string and the query string. The edit distance is employed in determining relevance of the document as part of result ranking by detecting near-matches of a whole query or part of the query. The edit distance evaluates how close the query string is to a given data stream that includes document information such as TAUC (title, anchor text, URL, clicks) information, etc. The architecture includes the index-time splitting of compound terms in the URL to allow the more effective discovery of query terms. Additionally, index-time filtering of anchor text is utilized to find the top N anchors of one or more of the document results. The TAUC information can be input to a neural network (e.g., 2-layer) to improve relevance metrics for ranking the search results.
公开/授权文献
信息查询