发明名称 Method of obtaining data samples from a data stream and of estimating the sortedness of the data stream based on the samples
摘要 Disclosed is a method of scanning a data stream in a single pass to obtain uniform data samples from selected intervals. The method comprises randomly selecting elements from the stream for storage in one or more data buckets and, then, randomly selecting multiple samples from the bucket(s). Each sample is associated with a specified interval immediately prior to a selected point in time. There is a balance of probabilities between the selection of elements stored in the bucket and the selection of elements included in the samples so that elements scanned during the specified interval are included in the sample with equal probability. Samples can then be used to estimate the degree of sortedness of the stream, based on counting how many elements in the sequence are the rightmost point of an interval such that majority of the interval's elements are inverted with respect to the interval's rightmost element.
申请公布号 US7797326(B2) 申请公布日期 2010.09.14
申请号 US20060405994 申请日期 2006.04.18
申请人 INTERNATIONAL BUSINESS MACHINES CORPORATION 发明人 GOPALAN PARIKSHIT;KRAUTHGAMER ROBERT;THATHACHAR JAYRAM S.
分类号 G06F7/00;G06F17/30 主分类号 G06F7/00
代理机构 代理人
主权项
地址