摘要 |
<p><P>PROBLEM TO BE SOLVED: To provide a method for detecting similar objects in a collection of such objects. <P>SOLUTION: The method modifies a previous method in such a way that per-object memory requirements are reduced while false detections are avoided approximately as well as in the previous method. The modification includes (i) combining k samples of features into s supersamples, the value of k being reduced from the corresponding value used in the previous method; (ii) recording each supersample to b bits of precision, the value of b being reduced from the corresponding value used in the previous method; and (iii) requiring l matching supersamples in order to conclude that the two objects are sufficiently similar, the value of l being greater than the corresponding value required in the previous method. One application of the method is in association with a web search engine query service to determine clusters of query results that are look-alike documents (similar documents). <P>COPYRIGHT: (C)2006,JPO&NCIPI</p> |