Show simple item record

dc.contributor.authorDu, Qianzhou
dc.contributor.authorZhang, Xuan
dc.date.accessioned2015-05-13T01:34:58Z
dc.date.available2015-05-13T01:34:58Z
dc.date.issued2015-05-10
dc.identifier.urihttp://hdl.handle.net/10919/52254
dc.descriptionThis project explored how to apply Named Entity Recognition to large Twitter and web page datasets to extract useful entities such as people, organization, location, and date. In addition, this NER utility has been scaled to the MapReduce framework on the Hadoop cluster. A schema and software allow this to be integrated with IDEAL.en_US
dc.description.abstractThe term “Named Entity”, which was first introduced by Grishman and Sundheim, is widely used in Natural Language Processing (NLP). The researchers were focusing on the information extraction task, that is extracting structured information of company activities and defense related activities from unstructured text, such as newspaper articles. The essential part of “Named Entity” is to recognize information elements, such as location, person, organization, time, date, money, percent expression, etc. To identify these entities from unstructured text, some researchers called this sub-task of information extraction as “Named Entity Recognition” (NER). Now, NER technology has become mature and there are good tools to implement this task, such as the Stanford Named Entity Recognizer (SNER), Illinois Named Entity Tagger (INET), Alias-i LingPipe (LIPI), and OpenCalasi (OCWS). Each of these has some advantages and is designed for some special data. In this term project, our final goal is to build a NER module for the IDEAL project based on a particular NER tool, such as SNER, to apply NER to the Twitter and web pages data sets. This project report presents our work towards this goal, including literature review, requirements, algorithm, development plan, system architecture, implementation, user manual, and development manual. Further, results are given with regard to multiple collections, along with discussion and plans for the future.en_US
dc.description.sponsorshipNSF grant IIS - 1319578, III: Small: Integrated Digital Event Archiving and Library (IDEAL)en_US
dc.language.isoen_USen_US
dc.rightsAttribution-ShareAlike 3.0 United States*
dc.rights.urihttp://creativecommons.org/licenses/by-sa/3.0/us/*
dc.subjectNamed Entity Recognitionen_US
dc.subjectInformation Extractionen_US
dc.subjectInformation Retrievalen_US
dc.subjectMapReduceen_US
dc.subjectHadoopen_US
dc.titleNamed Entity Recognition for IDEALen_US
dc.typePresentationen_US
dc.typeSoftwareen_US
dc.typeTechnical reporten_US


Files in this item

Thumbnail
Thumbnail
Thumbnail
Thumbnail
Thumbnail
Thumbnail
Thumbnail

This item appears in the following Collection(s)

Show simple item record

Attribution-ShareAlike 3.0 United States
License: Attribution-ShareAlike 3.0 United States