Named Entity Recognition for IDEAL

Du, Qianzhou; Zhang, Xuan

Named Entity Recognition for IDEAL

dc.contributor.author	Du, Qianzhou	en
dc.contributor.author	Zhang, Xuan	en
dc.date.accessioned	2015-05-13T01:34:58Z	en
dc.date.available	2015-05-13T01:34:58Z	en
dc.date.issued	2015-05-10	en
dc.description	This project explored how to apply Named Entity Recognition to large Twitter and web page datasets to extract useful entities such as people, organization, location, and date. In addition, this NER utility has been scaled to the MapReduce framework on the Hadoop cluster. A schema and software allow this to be integrated with IDEAL.	en
dc.description.abstract	The term “Named Entity”, which was first introduced by Grishman and Sundheim, is widely used in Natural Language Processing (NLP). The researchers were focusing on the information extraction task, that is extracting structured information of company activities and defense related activities from unstructured text, such as newspaper articles. The essential part of “Named Entity” is to recognize information elements, such as location, person, organization, time, date, money, percent expression, etc. To identify these entities from unstructured text, some researchers called this sub-task of information extraction as “Named Entity Recognition” (NER). Now, NER technology has become mature and there are good tools to implement this task, such as the Stanford Named Entity Recognizer (SNER), Illinois Named Entity Tagger (INET), Alias-i LingPipe (LIPI), and OpenCalasi (OCWS). Each of these has some advantages and is designed for some special data. In this term project, our final goal is to build a NER module for the IDEAL project based on a particular NER tool, such as SNER, to apply NER to the Twitter and web pages data sets. This project report presents our work towards this goal, including literature review, requirements, algorithm, development plan, system architecture, implementation, user manual, and development manual. Further, results are given with regard to multiple collections, along with discussion and plans for the future.	en
dc.description.sponsorship	NSF grant IIS - 1319578, III: Small: Integrated Digital Event Archiving and Library (IDEAL)	en
dc.identifier.uri	http://hdl.handle.net/10919/52254	en
dc.language.iso	en_US	en
dc.rights	Creative Commons Attribution-ShareAlike 3.0 United States	en
dc.rights.uri	http://creativecommons.org/licenses/by-sa/3.0/us/	en
dc.subject	Named Entity Recognition	en
dc.subject	Information Extraction	en
dc.subject	Information Retrieval	en
dc.subject	MapReduce	en
dc.subject	Hadoop	en
dc.title	Named Entity Recognition for IDEAL	en
dc.type	Presentation	en
dc.type	Software	en
dc.type	Technical report	en