Table Understanding for Information Retrieval

Pande, Ashwini K.

Table Understanding for Information Retrieval

Files

AshwiniPandeTableIR.pdf (869.98 KB)

Downloads: 367

Date

2002-08-19

Authors

Pande, Ashwini K.

Publisher

Virginia Tech

Abstract

This thesis proposes a novel approach for finding tables in text files containing a mixture of unstructured and structured text. Tables may be arbitrarily complex because the data in the tables may themselves be tables and because the grouping of data elements displayed in a table may be very complex. Although investigators have proposed competence models to explain the structure of tables, there are no computationally feasible performance models for detecting and parsing general structures in real data. Our emphasis is placed on the investigation of a new statistical procedure for detecting basic tables in plain text documents. The main task here is defining and testing this theory in the context of the Odessa Digital Library.

Keywords

Information retrieval, Statistical crosscorrelation, Odessa digital library, detection heuristics, Table detection

Persistent link

http://hdl.handle.net/10919/34820

Collections

Masters Theses

Full item page

Table Understanding for Information Retrieval

Files

TR Number

Date

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Description

Keywords

Citation

Persistent link

Collections