Korean-Chinese person name translation for cross language information retrieval

Yu Chun Wang, Yi Hsun Lee, Chu Cheng Lin, Richard Tzong Han Tsai, Wen Lian Hsu

研究成果: 會議貢獻類型會議論文同行評審

摘要

Named entity translation plays an important role in many applications, such as information retrieval and machine translation. In this paper, we focus on translating person names, the most common type of name entity in Korean-Chinese cross language information retrieval (KCIR). Unlike other languages, Chinese uses characters (ideographs), which makes person name translation difficult because one syllable may map to several Chinese characters. We propose an effective hybrid person name translation method to improve the performance of KCIR. First, we use Wikipedia as a translation tool based on the inter-language links between the Korean edition and the Chinese or English editions. Second, we adopt the Naver people search engine to find the query name's Chinese or English translation. Third, we extract Korean-English transliteration pairs from Google snippets, and then search for the English-Chinese transliteration in the database of Taiwan's Central News Agency or in Google. The performance of KCIR using our method is over five times better than that of a dictionary-based system. The mean average precision is 0.3490 and the average recall is 0.7534. The method can deal with Chinese, Japanese, Korean, as well as non-CJK person name translation from Korean to Chinese. Hence, it substantially improves the performance of KCIR.

原文???core.languages.en_GB???
頁面489-497
頁數9
出版狀態已出版 - 2007
事件21st Pacific Asia Conference on Language, Information and Computation, PACLIC 21 - Seoul, Korea, Republic of
持續時間: 1 11月 20073 11月 2007

???event.eventtypes.event.conference???

???event.eventtypes.event.conference???21st Pacific Asia Conference on Language, Information and Computation, PACLIC 21
國家/地區Korea, Republic of
城市Seoul
期間1/11/073/11/07

指紋

深入研究「Korean-Chinese person name translation for cross language information retrieval」主題。共同形成了獨特的指紋。

引用此