MetaEnhance: Metadata Quality Improvement for Electronic Theses and Dissertations of University Libraries
dc.contributor.author | Choudhury, Muntabir Hasan | en |
dc.contributor.author | Salsabil, Lamia | en |
dc.contributor.author | Jayanetti, Himarsha R. | en |
dc.contributor.author | Wu, Jian | en |
dc.contributor.author | Ingram, William A. | en |
dc.contributor.author | Fox, Edward A. | en |
dc.date.accessioned | 2024-01-22T13:08:02Z | en |
dc.date.available | 2024-01-22T13:08:02Z | en |
dc.date.issued | 2023 | en |
dc.description.abstract | Metadata quality is crucial for discovering digital objects through digital library (DL) interfaces. However, due to various reasons, the metadata of digital objects often exhibits incomplete, inconsistent, and incorrect values. We investigate methods to automatically detect, correct, and canonicalize scholarly metadata, using seven key fields of electronic theses and dissertations (ETDs) as a case study. We propose MetaEnhance, a framework that utilizes state-of-the-art artificial intelligence (AI) methods to improve the quality of these fields. To evaluate MetaEnhance, we compiled a metadata quality evaluation benchmark containing 500 ETDs, by combining subsets sampled using multiple criteria. We evaluated MetaEnhance against this benchmark and found that the proposed methods achieved nearly perfect F1-scores in detecting errors and F1-scores ranging from 0.85 to 1.00 for correcting five of seven key metadata fields. The codes and data are publicly available on GitHub11https://github.com/lamps-lab/ETDMiner/tree/master/metadata-correction. | en |
dc.description.version | Submitted version | en |
dc.format.extent | Pages 61-65 | en |
dc.format.extent | 5 page(s) | en |
dc.format.mimetype | application/pdf | en |
dc.identifier.doi | https://doi.org/10.1109/JCDL57899.2023.00019 | en |
dc.identifier.eissn | 2575-8152 | en |
dc.identifier.isbn | 9798350399318 | en |
dc.identifier.issn | 2575-7865 | en |
dc.identifier.orcid | Ingram, William [0000-0002-8307-8844] | en |
dc.identifier.orcid | Fox, Edward [0000-0003-1447-6870] | en |
dc.identifier.uri | https://hdl.handle.net/10919/117434 | en |
dc.identifier.volume | 2023-June | en |
dc.language.iso | en | en |
dc.publisher | ACM | en |
dc.rights | In Copyright | en |
dc.rights.uri | http://rightsstatements.org/vocab/InC/1.0/ | en |
dc.subject | Digital Libraries | en |
dc.subject | Scholarly Big Data | en |
dc.subject | ETD | en |
dc.subject | Metadata Quality | en |
dc.subject | Artificial Intelligence | en |
dc.title | MetaEnhance: Metadata Quality Improvement for Electronic Theses and Dissertations of University Libraries | en |
dc.title.serial | 2023 ACM/IEEE JOINT CONFERENCE ON DIGITAL LIBRARIES, JCDL | en |
dc.type | Conference proceeding | en |
dc.type.dcmitype | Text | en |
dc.type.other | Proceedings Paper | en |
dc.type.other | Book in series | en |
pubs.finish-date | 2023-06-30 | en |
pubs.organisational-group | /Virginia Tech | en |
pubs.organisational-group | /Virginia Tech/Engineering | en |
pubs.organisational-group | /Virginia Tech/Engineering/Computer Science | en |
pubs.organisational-group | /Virginia Tech/Library | en |
pubs.organisational-group | /Virginia Tech/All T&R Faculty | en |
pubs.organisational-group | /Virginia Tech/Engineering/COE T&R Faculty | en |
pubs.organisational-group | /Virginia Tech/Library/Library assessment administrators | en |
pubs.organisational-group | /Virginia Tech/Library/Dean's office | en |
pubs.organisational-group | /Virginia Tech/Library/Information Technology | en |
pubs.organisational-group | /Virginia Tech/Graduate students | en |
pubs.organisational-group | /Virginia Tech/Graduate students/Doctoral students | en |
pubs.start-date | 2023-06-26 | en |