摘要 |
A method for estimating artist ambiguity in a dataset is performed at a device with a processor and memory storing instructions for execution by the processor. The method includes applying a statistical classifier to a first dataset including a plurality of media items, wherein each media item is associated with one of a plurality of artist identifiers, each artist identifier identifies a real world artist, and the statistical classifier calculates a respective probability that each respective artist identifier is associated with media items from two or more different real world artists based on a respective feature vector corresponding to the respective artist identifier. The method further includes providing a report of the first dataset, including the calculated probabilities, to a user of the electronic device. Each respective feature vector includes a plurality of features that indicate likelihood of artist ambiguity. |