发明名称 System and method for interactive classification and analysis of data
摘要 A system, method, and computer program product for interactively classifying and analyzing data is particularly applicable to classification and analysis of textual data. It is particularly useful in identification of helpdesk inquiry and problem categories amenable to automated fulfillment or solution. A dictionary is generated based on a frequency of occurrence of words in a document set. A count of occurrences of each word in the dictionary within each document in the document set is generated. The set of documents is partitioned into a plurality of clusters. A name, a centroid, a cohesion score, and a distinctness score are generated for each cluster and displayed in a table. The documents contained in the clusters sorted based on their similarity to other documents in the cluster. The similarity may be determined by calculating the distance of the document to the centroid of the cluster and the documents may be sorted in order of ascending or descending distance of the document to the centroid of the cluster. Editing input may be received from a user and the displayed table modified based on the received editing input-clusters may be split or deleted. The helpdesk application area is only one of many areas to which the present invention may be advantageously applied. One of ordinary skill in the art would recognize that any set of text documents may be classified and subsequently analyzed using the present invention.
申请公布号 US6424971(B1) 申请公布日期 2002.07.23
申请号 US19990429650 申请日期 1999.10.29
申请人 INTERNATIONAL BUSINESS MACHINES CORPORATION 发明人 KREULEN JEFFREY THOMAS;MODHA DHARMENDRA SHANTILAL;SPANGLER WILLIAM SCOTT;STRONG, JR. HOVEY RAYMOND
分类号 G06F17/30;(IPC1-7):G06F17/30 主分类号 G06F17/30
代理机构 代理人
主权项
地址