Repository logo
Article

Named-entity recognition for Hindi language using context pattern-based maximum entropy

creativeworkseries.issn1508-2806
dc.contributor.authorJain, Arti
dc.contributor.authorYadav, Divakar
dc.contributor.authorArora, Anuja
dc.contributor.authorTayal, Devendra K.
dc.date.available2025-06-20T04:40:21Z
dc.date.issued2022
dc.descriptionBibliogr. s. 101-112.
dc.description.abstractThis paper describes a named-entity-recognition (NER) system for the Hindi language that uses two methodologies: an existing baseline maximum entropy-based named-entity (BL-MENE) model, and the proposed context pattern-based MENE (CP-MENE) framework. BL-MENE utilizes several baseline features for the NER task but suffers from inaccurate named-entity (NE) boundary detection, misclassification errors, and the partial recognition of NEs due to certain missing essentials. However, the CP-MENE-based NER task incorporates extensive features and patterns that are set to overcome these problems. In fact, CP-MENE’s features include right-boundary, left-boundary, part-of-speech, synonym, gazetteer and relative pronoun features. CP-MENE formulates a kind of recursive relationship for extracting highly ranked NE patterns that are generated through regular expressions via Python@ code. Since the web content of the Hindi language is arising nowadays (especially in health care applications), this work is conducted on the Hindi health data (HHD) corpus (which is readily available from the Kaggle dataset). Our experiments were conducted on four NE categories, namely, Person (PER), Disease (DIS), Consumable (CNS), and Symptom (SMP).en
dc.description.placeOfPublicationKraków
dc.description.versionwersja wydawnicza
dc.identifier.doihttps://doi.org/10.7494/csci.2022.23.1.3977
dc.identifier.eissn2300-7036
dc.identifier.issn1508-2806
dc.identifier.urihttps://repo.agh.edu.pl/handle/AGH/113299
dc.language.isoeng
dc.publisherWydawnictwa AGH
dc.relation.ispartofComputer Science
dc.rightsAttribution 4.0 International
dc.rights.accessotwarty dostęp
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/legalcode
dc.subjectcontext patternsen
dc.subjectgazetteer listsen
dc.subjectHindi languageen
dc.subjectKaggle dataseten
dc.subjectmaximum entropyen
dc.subjectnamed-entity recognitionen
dc.subjectfeature extensionen
dc.titleNamed-entity recognition for Hindi language using context pattern-based maximum entropyen
dc.title.relatedComputer Scienceen
dc.typeartykuł
dspace.entity.typePublication
publicationissue.issueNumberNo. 1
publicationissue.paginationpp. 81-115
publicationvolume.volumeNumberVol. 23
relation.isJournalIssueOfPublicationf31834f3-1961-48d0-8f61-017ec7fec754
relation.isJournalIssueOfPublication.latestForDiscoveryf31834f3-1961-48d0-8f61-017ec7fec754
relation.isJournalOfPublication020291ee-249b-4dcf-98a3-276a2f7981aa

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
csci.2022.23.1.81.pdf
Size:
277.07 KB
Format:
Adobe Portable Document Format