发明名称 SEARCH RESULTS RANKING USING EDITING DISTANCE AND DOCUMENT INFORMATION
摘要 Architecture for extracting document information from documents received as search results based on a query string, and computing an edit distance between the data string and the query string. The edit distance is employed in determining relevance of the document as part of result ranking by detecting near-matches of a whole query or part of the query. The edit distance evaluates how close the query string is to a given data stream that includes document information such as TAUC (title, anchor text, URL, clicks) information, etc. The architecture includes the index-time splitting of compound terms in the URL to allow the more effective discovery of query terms. Additionally, index-time filtering of anchor text is utilized to find the top N anchors of one or more of the document results. The TAUC information can be input to a neural network (e.g., 2-layer) to improve relevance metrics for ranking the search results.
申请公布号 ZA201006093(B) 申请公布日期 2011.10.26
申请号 ZA20100006093 申请日期 2010.08.26
申请人 MICROSOFT CORPORATION 发明人 TANKOVICH VLADIMIR;LI HANG;MEYERZON DMITRIY;XU JUN
分类号 G06F 主分类号 G06F
代理机构 代理人
主权项
地址