摘要 |
A method and apparatus of processing data is disclosed for recognizing unknown characters of a known character set, some of the characters having diacritical marks. The method includes the steps of storing the image data of an unknown character which may contain a diacritical mark. From the stored image data a predetermined localized area of data is extracted that corresponds to the expected location of the diacritical mark. The extracted diacritical mark image data and at least a portion of the stored image data of the unknown character are examined to recognize the character and any diacritical mark associated therewith. Also disclosed are video preprocessing techniques for segmenting the characters using profiles thereof, inclusive-bit-coding to separate characters based upon differences in size, justification of the extracted diacritical mark image data, unique encoding of the recognition results, and post-processing verification for characters including diacritical marks.
|