发明名称 Extracting sentence translations from translated documents
摘要 A system extracts translations from translated texts, such as sentence translations from translated versions of documents. A first and a second text are accessed and divided into a plurality of textual elements. From these textual elements, a sequence of pairs of text portions is formed, and a pair score is calculated for each pair, using weighted features. Then, an alignment score of the sequence is calculated using the pair scores, and the sequence is systematically varied to identify a sequence that optimizes the alignment score. The invention allows for fast, reliable and robust alignment of sentences within large translated documents. Further, it allows to exploit a broad variety of existing knowledge sources in a flexible way, without performance penalty. Further, a general implementation of dynamic programming search with online memory allocation and garbage collection allows for treating very long documents with limited memory footprint. <IMAGE>
申请公布号 EP1227409(A2) 申请公布日期 2002.07.31
申请号 EP20010130198 申请日期 2001.12.19
申请人 XEROX CORPORATION 发明人 EISELE, ANDREAS
分类号 G06F17/27;G06F17/28 主分类号 G06F17/27
代理机构 代理人
主权项
地址