发明名称 Method and apparatus for inferring the topical content of a document based upon its lexical content without supervision
摘要 An iterative method of determining the topical content of a document using a computer. The processing unit of the computer determines the topical content of documents presented to it in machine readable form using information stored in computer memory. That information includes word-clusters, a lexicon, and association strength values. The processing unit beings by generating an observed feature vector for the document being characterized, which indicates which of the words of the lexicon appear in the document. Afterward, the processing unit makes an initial prediction of the topical content of the document in the form of a topic belief vector. The processing unit uses the topic belief vector and the association strength values to predict which words of the lexicon should appear in the document. This prediction is represented via a predicted feature vector. The predicted feature vector is then compared to the observed feature vector to measure how well the topic belief vector models the topical content of the document. If the topic belief vector adequately model the topical content of the document, then the processing unit's task is complete. On the other hand, if the topic belief vector does not adequately model the topical content of the document, then the processing unit determines how the topic belief vector should be modified to improve the prediction of modeling of the topical content.
申请公布号 US5659766(A) 申请公布日期 1997.08.19
申请号 US19940307221 申请日期 1994.09.16
申请人 XEROX CORPORATION 发明人 SAUND, ERIC;HEARST, MARTI A.
分类号 G06F17/30;(IPC1-7):G06F17/28 主分类号 G06F17/30
代理机构 代理人
主权项
地址