Show simple item record

dc.contributor.authorMehta, Sneha
dc.contributor.authorVinayagam, Radha Krishnan
dc.descriptionThis submission includes the project report, final presentation, LDA code, test datasets and its results. In the compressed folder, "", we have included the LDA Scala source code (lda_v1.scala) for processing Tweets and a JAR file for web page analysis. The compressed folder, "TopicAnalysis-TestData&" contains cleaned Tweet collections and web pages from the Obamacare collection. In the same folder, we have also included the topic results for each collection and a PDF file to interpret the collection IDs.en_US
dc.description.abstractThe IDEAL (Integrated Digital Event Archiving and Library) project aims to ingest tweets and web-based content from social media and the web and index it for retrieval. One of the required milestones for a graduate-level course CS5604 on Information Storage and Retrieval is to implement a state-of-the-art information retrieval and analysis system in support of the IDEAL project. The overall objective of this project is to build a robust Information Retrieval system on top of Solr, a general purpose open-source search engine. To enable the search and retrieval process we use various approaches including Latent Dirichlet Allocation, Named-Entity Recognition, Clustering, Classification, Social Network Analysis and Front-end interface for search. The project has been divided into various segments and our team has been assigned Topic Analysis. A topic in this context is a set of words that can be used to represent a document. The output of our team will be a well-defined set of topics that describe each document in the collections we have. The topics will facilitate a facet based search in the frontend search interface. This submission includes the project report, final presentation, LDA code, test datasets, and results. In the project report,we introduce the relevant background, design & implementation, and the requirements to make our part functional. The developer’s manual describes our approach in detail. Walk-through tutorials for related software packages have been included in the user’s manual. Finally, we also provide exhaustive results and detailed evaluation methodologies for the topic quality.en_US
dc.description.sponsorshipNSF IIS - 1319578: III: Small: Integrated Digital Event Archiving and Library (IDEAL)en_US
dc.rightsCC0 1.0 Universal*
dc.subjectTopic Analysisen_US
dc.subjectInformation Retrievalen_US
dc.titleTopic Analysis project in CS5604, Spring 2016: Extracting Topics from Tweets and Webpages for IDEALen_US
dc.typeTechnical reporten_US

Files in this item


This item appears in the following Collection(s)

Show simple item record

CC0 1.0 Universal
License: CC0 1.0 Universal