Automated Testing for Data-Intensive Scalable Computing

dc.contributor.authorHumayun, Ahmaden
dc.contributor.committeechairGulzar, Muhammad Alien
dc.contributor.committeememberTilevich, Elien
dc.contributor.committeememberRavindran, Binoyen
dc.contributor.committeememberKim, Miryungen
dc.contributor.committeememberWilliams, Daniel Johnen
dc.contributor.departmentComputer Science and#38; Applicationsen
dc.date.accessioned2026-08-06T08:00:30Zen
dc.date.available2026-08-06T08:00:30Zen
dc.date.issued2026-08-05en
dc.description.abstractData-Intensive Scalable Computing (DISC) systems such as Apache Spark, Hadoop, Flink, and Beam have become central to modern data processing. These systems enable developers to write applications as dataflows composed of user-defined functions operating over large and often unstructured datasets. However, testing such applications and the underlying frameworks remains a significant challenge. Conventional testing techniques struggle in this space due to the complexity of dataflow semantics, the scale and heterogeneity of input data, and the need to reason about operator interactions, schema variability, and program logic simultaneously. This thesis presents a set of techniques for extracting and leveraging rich, fine-grained properties that are unique to DISC workloads in order to inform automated input generation for effective testing across various levels of the data-centric software stack. The first thrust of this work focuses on generating inputs that can avoid trivial parsing errors and effectively exercise the deeper logic in DISC applications. By analyzing how code interacts with different parts of the input data, we test the code where it matters most instead of wasteful fuzzing cycles finding parsing issues. The second thrust explores how to create realistic and meaningful test data that reflects the structure and semantics of real-world inputs, while still achieving high coverage and fault detection. The final part of this work shifts focus to DISC frameworks, recognizing that applications are only as reliable as the systems they run on. By generating diverse dataflow programs, we systematically test internal components like optimizers, extending fuzzing to the entire DISC stack.en
dc.description.abstractgeneralModern data-intensive applications operate at extraordinary scale; historical sales analyses at companies like Amazon may require examining billions of records. To handle such volumes, practitioners rely on Data-Intensive Scalable Computing (DISC), a paradigm in which a program coordinates thousands of machines to process massive datasets in parallel. DISC frameworks, such as Apache Spark and Hadoop, provide the underlying execution machinery, allowing developers to express complex data transformations without managing the low- level details of distributed computing by hand. As DISC programs and their frameworks grow in complexity, automated testing techniques designed for conventional single-machine software are proving inadequate, creating an urgent need for new approaches tailored to the distributed setting. This thesis presents novel methods for automated testing across the full data processing stack. The first contribution proposes a technique for detecting faults in programs that combine and cross-reference multiple large datasets. The second develops methods that generate realistic-looking input data while simultaneously exposing program defects. The third introduces techniques to test the underlying infrastructure that coordinates and executes DISC computations, a layer that prior work has largely left untested.en
dc.description.degreeDoctor of Philosophyen
dc.format.mediumETDen
dc.identifier.othervt_gsexam:47427en
dc.identifier.urihttps://hdl.handle.net/10919/143697en
dc.language.isoenen
dc.publisherVirginia Techen
dc.rightsIn Copyrighten
dc.rights.urihttp://rightsstatements.org/vocab/InC/1.0/en
dc.subjectAutomated Testingen
dc.subjectFuzzingen
dc.subjectScalable Computingen
dc.titleAutomated Testing for Data-Intensive Scalable Computingen
dc.typeDissertationen
thesis.degree.disciplineComputer Science & Applicationsen
thesis.degree.grantorVirginia Polytechnic Institute and State Universityen
thesis.degree.leveldoctoralen
thesis.degree.nameDoctor of Philosophyen

Files

Original bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Humayun_A_D_2026.pdf
Size:
3.6 MB
Format:
Adobe Portable Document Format