摘要 |
A sequence of documents is delivered to an optical scanner in which each document is scanned to form a digital image representation of the content of the document. The image representation is automatically examined by data processing apparatus to select search words which meet predetermined criteria and by which the document can be subsequently located. The search words are stored in a non-volatile memory and the entire document content is stored in mass storage in image form. A font table is established and images of entered search words are constructed from the table. Unrecognized or imperfectly formed ambiguous characters are stored with the font table and are used in the construction of search words to ease or eliminate text editing or are stored with converted characters of search words. The approach can also be used to ease or eliminate editing of full text-converted data bases during input.
|