Named-entity recognition for Hindi language using context pattern-based maximum entropy

Szczegóły
Opis

Tytuł:: Named-entity recognition for Hindi language using context pattern-based maximum entropy
Autorzy:: Jain, Arti
Yadav, Divakar
Arora, Anuja
Tayal, Devendra K.
Powiązania:: https://bibliotekanauki.pl/articles/27312839.pdf
Data publikacji:: 2022
Wydawca:: Akademia Górniczo-Hutnicza im. Stanisława Staszica w Krakowie. Wydawnictwo AGH
Tematy:: context patterns
gazetteer lists
Hindi language
Kaggle dataset
maximum entropy
named-entity recognition
feature extension
Źródło:: Computer Science; 2022, 23 (1); 81--115
1508-2806
2300-7036
Język:: angielski
Prawa:: CC BY: Creative Commons Uznanie autorstwa 4.0
Dostawca treści:: Biblioteka Nauki
: Artykuł

Przejdź do źródła

This paper describes a named-entity-recognition (NER) system for the Hindi language that uses two methodologies: an existing baseline maximum entropy-based named-entity (BL-MENE) model, and the proposed context pattern-based MENE (CP-MENE) framework. BL-MENE utilizes several baseline features for the NER task but suffers from inaccurate named-entity (NE) boundary detection, misclassification errors, and the partial recognition of NEs due to certain missing essentials. However, the CP-MENE-based NER task incorporates extensive features and patterns that are set to overcome these problems. In fact, CP-MENE’s features include right-boundary, left-boundary, part-of-speech, synonym, gazetteer and relative pronoun features. CP-MENE formulates a kind of recursive relationship for extracting highly ranked NE patterns that are generated through regular expressions via Python@ code. Since the web content of the Hindi language is arising nowadays (especially in health care applications), this work is conducted on the Hindi health data (HHD) corpus (which is readily available from the Kaggle dataset). Our experiments were conducted on four NE categories; namely, Person (PER), Disease (DIS), Consumable (CNS), and Symptom (SMP).

Informacja

Powiązane pozycje