摘要 |
An incorrect hyperlink detecting apparatus which can detect a semantic inconsistency of a hyperlink with high accuracy is provided. An incorrect hyperlink detecting apparatus 10 includes a link source text extracting unit 12 for extracting a text from an HTML file 26 of a link source, a link destination text extracting unit 14 for extracting a text from the HTML file 26 of a link destination, a morpheme analysis unit 18 for dissolving the extracted texts into words, a weighting unit 18 for assigning a weightier every part of speech, a consistency rate calculating unit 20 for calculating a rate that the words of the link source are included in the words of the Sink destination as a consistency rate from the link source to the link destination and a rate that the words of the Sink destination are included in the words of the Sink source as a consistency rate from the link destination to the link source, degree of association calculating unit 22 for calculating a degree of association which indicates a probability of the hyperlink in response to both of the consistency rates, and a CSV output unit 24 for outputting the consistency rate and the degree of association in a CSV form.
|