发明名称 IDENTIFYING LANGUAGE AND CHARACTER SET OF DATA REPRESENTING TEXT
摘要 <p>The present invention provides a facility for identifying the unknown language of text represented by a series of data values in accordance with a character set that associates character glyphs with particular data values. The facility first generates a characterization that characterizes the series of data values in terms of the occurrence of particular data values on the series of data values. For each of a plurality of languages, the facility then retrieves a model that models the language in term of the statistical occurrence of particular data values in representative samples of text in that language. The facility then compares the retrieved models to the generated characterization of the series of data values, and identifies as the distinguished language the language whose model compares most favorably to the generated characterization of the series of data values.</p>
申请公布号 WO1999030252(A1) 申请公布日期 1999.06.17
申请号 US1998025814 申请日期 1998.12.04
申请人 发明人
分类号 主分类号
代理机构 代理人
主权项
地址