发明名称 PHRASE-BASED DETECTION OF DUPLICATE DOCUMENTS IN AN INFORMATION RETRIEVAL SYSTEM
摘要 An information retrieval system uses phrases to index, retrieve, organize and describe documents. Phrases are identified that predict the presence of other phrases in documents. Documents are the indexed according to their included phrases. Related phrases and phrase extensions are also identified. Phrases in a query are identified and used to retrieve and rank documents. Phrases are also used to cluster documents in the search results, create document descriptions, and eliminate duplicate documents from the search results, and from the index.
申请公布号 US2008306943(A1) 申请公布日期 2008.12.11
申请号 US20040900012 申请日期 2004.07.26
申请人 PATTERSON ANNA LYNN 发明人 PATTERSON ANNA LYNN
分类号 G06F7/00 主分类号 G06F7/00
代理机构 代理人
主权项
地址
您可能感兴趣的专利