摘要 |
A method and system for generating an audio thumbnail of an audio track in which a first content feature, such as singing, is detected as a characteristic of an audio track. A predetermined length of the detected portion of the audio track corresponding to the first content feature is extracted from the audio track. A highlight of the audio track, such as a portion of the audio track having a sudden increase in temporal energy within the audio track, is detected; and a portion of the audio track corresponding to the highlight is extracted from the audio track. The two extracted portions of the audio track are combined as a thumbnail of the audio track. |