On the Construction of Web NER Model Training Tool based on Distant Supervision

Chien Lung Chou, Chia Hui Chang, Yuan Hao Lin, Kuo Chun Chien

研究成果: 雜誌貢獻期刊論文同行評審

摘要

Named entity recognition (NER) is an important task in natural language understanding, as it extracts the key entities (person, organization, location, date, number, etc.) and objects (product, song, movie, activity name, etc.) mentioned in texts. However, existing natural language processing (NLP) tools (such as Stanford NER) recognize only general named entities or require annotated training examples and feature engineering for supervised model construction. Since not all languages or entities have public NER support, constructing a tool for NER model training is essential for low-resource language or entity information extraction. In this article, we study the problem of developing a tool to prepare training corpus from the Web with known seed entities for custom NER model training via distant supervision. The major challenge of automatic labeling lies in the long labeling time due to large corpus and seed entities as well as the concern to avoid false positive and false negative examples due to short and long seeds. To solve this problem, we adopt locality-sensitive hashing (LSH) for various length of seed entities. We conduct experiments on five types of entity recognition tasks, including Chinese person names, food names, locations, points of interest (POIs), and activity names to demonstrate the improvements with the proposed Web NER model construction tool. Because the training corpus is obtained by automatic labeling of the seed entity-related sentences, one could use either the entire corpus or the positive only sentences for model training. Based on the experimental results, we found the decision should depend on whether traditional linear chained conditional random fields (CRF) or deep neural network-based CRF is used for model training as well as the completeness of the provided seed list.

原文???core.languages.en_GB???
文章編號3422817
期刊ACM Transactions on Asian and Low-Resource Language Information Processing
19
發行號6
DOIs
出版狀態已出版 - 11月 2020

指紋

深入研究「On the Construction of Web NER Model Training Tool based on Distant Supervision」主題。共同形成了獨特的指紋。

引用此