Method for finding a reference token sequence in an original token string within a database of token strings using appended non-contiguous substrings,申请号US19950512794-传众专利搜索

发明名称	Method for finding a reference token sequence in an original token string within a database of token strings using appended non-contiguous substrings
摘要	This method non sequentially compares a reference sequence of tokens to an original sequence of tokens to determine subsequences of tokens which exactly or similarly match. The method has a novel approach for creating a large number of indexes by partitioning strings of tokens into substrings, appending non contiguous substrings together to form tuples, and creating indexes from the tuples. Indexes are created in this manner for both the original and reference strings. Techniques are also provided to approximately or exactly locate the substrings which where used to create the tuples and indexes from the original sequence of tokens. Original and reference indexes are compared and matches are tracked. Higher numbers of matches result in higher scores (votes) in a table and indicate a stronger similarity between the sequences on the the original and reference strings. Using this method, the degree of similarity can also be determined. The Method is useful when comparing a reference sequence of tokens to a large database of original strings of tokens. It has applications in the biological sciences (human genome mapping or analyzing proteins) and in image, speech, and music recognition.
申请公布号	US5577249(A)	申请公布日期	1996.11.19
申请号	US19950512794	申请日期	1995.08.08
申请人	INTERNATIONAL BUSINESS MACHINES CORPORATION	发明人	CALIFANO, ANDREA
分类号	G06F12/00;C12N15/09;G06F17/30;G06Q50/00;G10H1/00;G10L15/10;(IPC1-7):G06F17/30;G06F7/20	主分类号	G06F12/00
代理机构		代理人
主权项
地址