A Study of Computational Reproducibility using URLs Linking to Open Access Datasets and Software

dc.contributor.authorSalsabil, Lamiaen
dc.contributor.authorWu, Jianen
dc.contributor.authorChoudhury, Muntabiren
dc.contributor.authorIngram, William A.en
dc.contributor.authorFox, Edward A.en
dc.contributor.authorRajtmajer, Sarahen
dc.contributor.authorGiles, C. Leeen
dc.date.accessioned2022-10-19T16:53:38Zen
dc.date.available2022-10-19T16:53:38Zen
dc.date.issued2022-04-25en
dc.date.updated2022-10-19T15:07:58Zen
dc.description.abstractDatasets and software packages are considered important resources that can be used for replicating computational experiments. With the advocacy of Open Science and the growing interest of investigating reproducibility of scientific claims, including URLs linking to publicly available datasets and software packages has become an institutionalized part of research publications. In this preliminary study, we investigated the disciplinary dependency and chronological trends of including open access datasets and software (OADS) in electronic theses and dissertations (ETDs), based on a hybrid classifier called OADSClassifier, consisting of a heuristic and a supervised learning model. The classifier achieves the best F1 of 0.92.We found that the inclusion of OADS-URLs exhibited a strong disciplinary dependence and the fraction of ETDs containing OADS-URLs has been gradually increasing over the past 20 years.We developed and share a ground truth corpus consisting of 500 manually labeled sentences containing URLs from scientific papers. The dataset and source code are available at https://github.com/lamps-lab/oadsclassifier.en
dc.description.versionPublished versionen
dc.format.mimetypeapplication/pdfen
dc.identifier.doihttps://doi.org/10.1145/3487553.3524658en
dc.identifier.urihttp://hdl.handle.net/10919/112210en
dc.language.isoenen
dc.publisherACMen
dc.rightsIn Copyrighten
dc.rights.holderThe author(s)en
dc.rights.urihttp://rightsstatements.org/vocab/InC/1.0/en
dc.titleA Study of Computational Reproducibility using URLs Linking to Open Access Datasets and Softwareen
dc.typeArticle - Refereeden
dc.type.dcmitypeTexten

Files

Original bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
3487553.3524658.pdf
Size:
637.75 KB
Format:
Adobe Portable Document Format
Description:
Published version
License bundle
Now showing 1 - 1 of 1
Name:
license.txt
Size:
0 B
Format:
Item-specific license agreed upon to submission
Description: