Automated Testing for Data-Intensive Scalable Computing
| dc.contributor.author | Humayun, Ahmad | en |
| dc.contributor.committeechair | Gulzar, Muhammad Ali | en |
| dc.contributor.committeemember | Tilevich, Eli | en |
| dc.contributor.committeemember | Ravindran, Binoy | en |
| dc.contributor.committeemember | Kim, Miryung | en |
| dc.contributor.committeemember | Williams, Daniel John | en |
| dc.contributor.department | Computer Science and#38; Applications | en |
| dc.date.accessioned | 2026-08-06T08:00:30Z | en |
| dc.date.available | 2026-08-06T08:00:30Z | en |
| dc.date.issued | 2026-08-05 | en |
| dc.description.abstract | Data-Intensive Scalable Computing (DISC) systems such as Apache Spark, Hadoop, Flink, and Beam have become central to modern data processing. These systems enable developers to write applications as dataflows composed of user-defined functions operating over large and often unstructured datasets. However, testing such applications and the underlying frameworks remains a significant challenge. Conventional testing techniques struggle in this space due to the complexity of dataflow semantics, the scale and heterogeneity of input data, and the need to reason about operator interactions, schema variability, and program logic simultaneously. This thesis presents a set of techniques for extracting and leveraging rich, fine-grained properties that are unique to DISC workloads in order to inform automated input generation for effective testing across various levels of the data-centric software stack. The first thrust of this work focuses on generating inputs that can avoid trivial parsing errors and effectively exercise the deeper logic in DISC applications. By analyzing how code interacts with different parts of the input data, we test the code where it matters most instead of wasteful fuzzing cycles finding parsing issues. The second thrust explores how to create realistic and meaningful test data that reflects the structure and semantics of real-world inputs, while still achieving high coverage and fault detection. The final part of this work shifts focus to DISC frameworks, recognizing that applications are only as reliable as the systems they run on. By generating diverse dataflow programs, we systematically test internal components like optimizers, extending fuzzing to the entire DISC stack. | en |
| dc.description.abstractgeneral | Modern data-intensive applications operate at extraordinary scale; historical sales analyses at companies like Amazon may require examining billions of records. To handle such volumes, practitioners rely on Data-Intensive Scalable Computing (DISC), a paradigm in which a program coordinates thousands of machines to process massive datasets in parallel. DISC frameworks, such as Apache Spark and Hadoop, provide the underlying execution machinery, allowing developers to express complex data transformations without managing the low- level details of distributed computing by hand. As DISC programs and their frameworks grow in complexity, automated testing techniques designed for conventional single-machine software are proving inadequate, creating an urgent need for new approaches tailored to the distributed setting. This thesis presents novel methods for automated testing across the full data processing stack. The first contribution proposes a technique for detecting faults in programs that combine and cross-reference multiple large datasets. The second develops methods that generate realistic-looking input data while simultaneously exposing program defects. The third introduces techniques to test the underlying infrastructure that coordinates and executes DISC computations, a layer that prior work has largely left untested. | en |
| dc.description.degree | Doctor of Philosophy | en |
| dc.format.medium | ETD | en |
| dc.identifier.other | vt_gsexam:47427 | en |
| dc.identifier.uri | https://hdl.handle.net/10919/143697 | en |
| dc.language.iso | en | en |
| dc.publisher | Virginia Tech | en |
| dc.rights | In Copyright | en |
| dc.rights.uri | http://rightsstatements.org/vocab/InC/1.0/ | en |
| dc.subject | Automated Testing | en |
| dc.subject | Fuzzing | en |
| dc.subject | Scalable Computing | en |
| dc.title | Automated Testing for Data-Intensive Scalable Computing | en |
| dc.type | Dissertation | en |
| thesis.degree.discipline | Computer Science & Applications | en |
| thesis.degree.grantor | Virginia Polytechnic Institute and State University | en |
| thesis.degree.level | doctoral | en |
| thesis.degree.name | Doctor of Philosophy | en |
Files
Original bundle
1 - 1 of 1