Doctoral Dissertations

Permanent URI for this collection

Browse

Recent Submissions

Now showing 1 - 20 of 18538
  • Quantum Online Learning
    Zheng, Alice (Virginia Tech, 2026-08-12)
    Online learning is a framework for interfacing with environments whose structure is revealed through interactions; quantum information theory is a means of representing data and its processing consistent with modern understanding of physics. Their union is symbiotic---sequential processing is light on precious quantum resources, and quantum algorithms can surpass known bounds of classical algorithms---their combination being an area of independent interest. This thesis aims to exemplify the above point via five case studies on tomography, change detection, and function evaluation. In the process of deriving efficient tomography procedures, prior work has shown that quantum states can be learned in online adversarial environments. We extend this notion to subsets of positive semidefinite operators, including general quantum objects such as states, measurements, channels, strategies, co-strategies, and Gram matrices. We tailor a regularized follow-the-leader algorithm and achieve sublinear regret. As a byproduct, we show a generalization of Pinsker's inequality that accounts for differences in trace, its proof method more broadly applicable to general divergences. We further extend online learnability to objects in arbitrary physical theories. We derive mappings between Euclidean Jordan algebras and generalized probabilistic theories, using methods for the former to show that states in the latter are learnable. Specifically, we develop a projective version of the symmetric cone multiplicative weights update algorithm that achieves sublinear regret when learning states of generalized probabilistic theories. We consider the task of detecting changes in sequences of unknown quantum states with a constraint of no false positives, developing efficient and optimal online quantum algorithms. These include online algorithms with small amounts of quantum memory based on the swap test, and optimal algorithms projecting onto the symmetric subspace. We find circuit constructions for these algorithms, analyze their complexity, and show their equivalence under certain memory limitations. The proposed algorithms are applicable to quality control and anomaly detection in quantum devices, and may be of independent interest for identity testing and optimal cloning. We elaborate on the quality control aspect of changepoint methods by developing procedures for aiding the calibration of near-term quantum devices. These consist of algorithms for certification and changepoint detection in Hamiltonian dynamics, which enable continuous monitoring and trigger recalibration procedures. Using a cumulative sum procedure, we attain an asymptotically-optimal scaling of the average times to detection versus false positive, dependent primarily on Hamiltonian norm bounds. Lastly, we quantify the effect of memory limits on Boolean function evaluation of quantum sequences---a task applicable to distributed sensing and quantum reading of classical memories, among others. We show that minimum-error function evaluation in the memoryless regime is equivalent to the so-called pretty good measurement and hence shares its performance guarantee. Additionally, we characterize the precise set of functions on which this strategy performs optimally---affine functions. The above studies reveal insights about learnability as a universal property and the tradeoffs between performance and resource requirements in online settings, informing the choice of quantum settings best-suited for online learning. We conclude by summarizing these takeaways, along with listing remaining open questions.
  • Exploration of Dynamic Structure-Process-Property Relationships in Vitrimer-Like Materials and Multimodal Polymer-Clay Composites
    Luster, Larry E. (Virginia Tech, 2026-08-12)
    When designing and modeling composite materials, it is common to make simplifying assumptions regarding intermolecular interactions so that the composite can be considered as a homogeneous material with predictable behavior. This approach is acceptable for traditional composite processing and in any case where the components of the composite have uniform, static structures after they are incorporated into the bulk material; however, real materials for engineering applications rarely fit these idealized models and the components have specific, non-negligible interactions and transient topologies. This body of work examines two such materials: vitrimeric thermoplastic polymers and polymers reinforced with layered aluminosilicates; specifically, we 1) attempt to deconvolute contributions of bond exchange kinetics and segmental relaxation in a covalent adaptive network comprised of a poly(methyl-methacrylate)-poly(hydroxy-ethyl-methacrylate) copolymer (PMMA-PHEMA) crosslinked with dynamic aromatic disulfide bonds and 2) examine the role of in-situ dehydration on intercalation and exfoliation of montmorillonite agglomerates into nanoplatelets during melt-extrusion of a polyethylene terephthalate glycol (PETG)-montmorillonite-zeolite composite. In both studies, we challenge the fundamental assumptions used to simplify kinetic, thermodynamic, and transport properties of the material and utilize bulk rheological measurements to extrapolate mechanistic understandings of the material behavior that can be exploited in process design to yield desirable properties and morphologies in the end-use material.
  • Exploring the Potential Safety Impact of Automated Driving Systems Using Naturalistic Data
    Herbers, Eileen Mary (Virginia Tech, 2026-08-12)
    Automated driving systems (ADS) have the possibility to remove humans from the primary driving task, which has the potential to eliminate all crashes due to human driver error. However, a large component preventing the wide-scale adoption of ADS is the ability to confidently determine when they are safe enough to deploy and achieve the associated societal acceptance. The question of "how safe is safe enough?" has consumed the industry for quite some time. Determining the threshold of "safe enough," however, has proven to be quite a complex problem. A common perspective suggests that as long as ADS are safer than human drivers, then they are safe enough to deploy. However, this perspective introduces several challenges, including determining an appropriate human driver baseline (some drivers are safer than others), defining the conditions and scale of comparison, and considering whether society is willing to accept current roadway risks, which result in over forty thousand fatalities each year [1]. In these assessments, it is relevant to note that most driving is routine. Thus, modest ADS deployments are rarely exposed to the complex scenarios that the 233 million licensed human drivers encounter across the United States [2]. In these rare situations, human drivers might outperform ADS unless we identify methods to characterize these events and better understand what we are expecting robotic drivers to achieve. Naturalistic driving data, which represents large-scale in situ data collections of everyday driving, are used within this dissertation to identify and validate events and scenarios that might challenge the capabilities of ADS, indicating areas where further development and validation may be required to achieve acceptable levels of safety. Specifically, this dissertation analyzes nearly three thousand safety-critical events (SCEs) that involve crashes and near-crashes in a variety of scenarios. Specific events are selected to evaluate conditions under which ADS may have difficulty navigating the situation correctly. The analysis focuses on surprise events, including scenarios in which line of sight (LOS) perception systems are obstructed and constrain system response, as well as events involving unpredictable behavior by other road users in which the subject driver is not at fault. This analysis suggests that ADS may not perform as expected in blind turns and hills, mixed-speed traffic, lane-change events with other vehicles around, scenarios with significant occlusion at high speeds, and in scenarios in which pedestrians are visually occluded. By investigating SCE scenarios that are believed to be challenging for ADS to navigate, this research shows that using a small set of naturalistic data has the potential to convey important information to wide-scale ADS deployment that simulation or closed-track testing based on contrived scenarios simply cannot achieve. Near-crash and crash-relevant events are especially crucial for fully understanding the complex driving task and should be studied further to completely assess ADS safety. Additionally, human drivers are generally good at performing evasive maneuvers that require a complex understanding of the surrounding environment. Such near-crash situations necessitate an intricate sequence of perception and response for ADS, which may not be fully understood based on data from the limited ADS deployments performed to date. This dissertation aims to inform selection of ADS testing scenarios, catalogue edge cases which challenge ADS and limit their ability to provide safe transport, and to identify the opportunity to improve ADS performance in such situations through inclusion Vehicle-to-Everything (V2X) communications which enable perception beyond LOS sensing.
  • Robust Functional-Input Hypothesis Testing and Multi-Task Regression on Mixed-Type Graphs for High-Dimensional, Complex-Structured Data
    Chen, Mengkun (Virginia Tech, 2026-08-10)
    High-dimensional, complex-structured data have become increasingly prevalent across modern scientific disciplines such as social sciences, omics, and medical imaging. As such data continue to proliferate, developing advanced statistical methodologies for scalable estimation and inference is essential for extracting meaningful insights and driving scientific discovery. This dissertation introduces three robust, flexible, and computationally efficient nonparametric methods for functional data analysis: (1) a flexible test for detecting unknown functional departures under generalized functional regression for biomedical group discrimination; (2) a robust functional-input kernel-machine-based test for identifying brain networks associated with health outcomes; and (3) a joint functional regression on mixed-type functional graphs for dynamic network modeling in brain imaging data. For the first method, we introduce a hybrid Frequentist-Bayesian hypothesis testing procedure that combines a Bayes factor and a score-type statistic to detect nonlinear and nonparametric effects of functional predictors on scalar responses. This approach avoids restrictive parametric assumptions and explicit likelihood estimation, offering a practical and flexible tool for biomedical group discrimination. For the second method, we develop a kernel-machine-based quantile regression test to identify associations between high-dimensional, correlated functional predictors and heavy-tailed or skewed responses. This robust approach accommodates complex interactions and successfully identifies autism spectrum disorder (ASD)-related brain networks supported by neuroscience evidence. For the third method, we propose a unified joint modeling framework that simultaneously selects functional predictors embedded in a time-varying functional graph and estimates dynamic network structures without prior information. The model combines Bayesian hierarchical modeling by developing a computationally efficient marginal expected integrated penalized likelihood maximization (ME-IPLM) algorithm, enhancing variable and graph selection accuracy, interpretability, and prediction performance. Together, these three projects provide scalable, interpretable, and theoretically supported frameworks for analyzing high-dimensional functional data, advancing group discrimination, association detection, and dynamic network interpretation across diverse scientific domains.
  • Mechanisms of Nutrient Decline in Soybean under Elevated CO₂: A Physiological and Transcriptomic Investigation
    Kaur, Ravneet (Virginia Tech, 2026-08-10)
    Rising atmospheric carbon dioxide (CO₂) can stimulate photosynthesis, biomass production, and yield in C₃ crops, but these responses are often accompanied by lower seed mineral concentrations. The physiological and molecular processes underlying this decline remain poorly resolved in soybean (Glycine max L. Merr.), an important source of protein and minerals. This dissertation integrated physiological measurements, ionomic analysis, transcriptomics, and canopy-level seed sampling to determine how carbon assimilation, water use, nutrient uptake and allocation, gene expression, and canopy position contribute to soybean seed nutrient responses under elevated CO₂ (eCO₂). Soybean cultivars with contrasting yield, nutrient, and water-use responses were grown under ambient and elevated CO₂ in open-top or controlled-environment chambers. Across experiments, eCO₂ increased carbon assimilation, aboveground biomass, seed number, and seed yield in responsive cultivars, while root biomass and nutrient uptake did not increase proportionally with reproductive growth. Mature seed concentrations of several macro- and micronutrients declined, indicating that carbon-driven seed production exceeded the capacity of plants to acquire and allocate minerals to developing seeds. Transcriptomic analysis of leaves, roots, and seeds at the beginning seed stage showed that organ identity accounted for most variation in gene expression, with roots exhibiting the largest response to eCO₂. Leaf expression patterns were associated with photosynthesis, carbohydrate metabolism, and ion transport. Root responses were associated with cell wall modification, oxidative functions, hormone metabolism, and nutrient transport, while seed responses were associated with carbohydrate metabolism and cell wall modification. Most differentially expressed genes were restricted to individual cultivar-by-organ combinations, showing that similar seed mineral outcomes were not preceded by one shared transcriptional response. Instead, the combined physiological and molecular data indicated a disconnect between carbon-driven growth and proportional nutrient uptake and allocation. This interpretation was further evaluated using cultivars selected for contrasting transpiration and yield responses. Elevated CO₂ reduced transpiration by 8.4% and stomatal conductance by 37.5%, but seed nutrient accumulation followed cultivar yield response rather than water-use phenotype. High-yielding cultivars accumulated more total minerals in seeds, although this increase did not keep pace with seed production, whereas cultivars with little or negative yield responses accumulated fewer seed minerals. Canopy position also strongly influenced nutrient distribution, with middle-canopy seeds contributing the greatest seed mass and mineral accumulation regardless of CO₂ treatment. Together, these findings show that soybean seed nutrient decline under eCO₂ is driven primarily by carbon assimilation and reproductive growth outpacing nutrient acquisition and allocation, while reduced transpiration provides a less consistent explanation. The cultivar- and organ-specific responses identified here demonstrate that similar declines in seed mineral concentration can arise through different physiological and transcriptional relationships. This work provides a framework for identifying soybean traits and cultivars that maintain seed nutritional quality under future atmospheric conditions.
  • Defending the Throne: Leader Narcissism, Status Threat, and the Self-Regulatory Path to Abusive Supervision Intentions
    Conger, Joseph Zachary (Virginia Tech, 2026-08-10)
    This dissertation examined the relationships between unidimensional and tripartite narcissism, four mediating cognitive self-regulatory responses, and abusive supervision intentions. I expected unidimensional narcissism to have positive associations with the selected self-regulatory responses and abusive supervision intentions. I recruited 202 full-time workers with supervisory experience to partake in a between-person, paper-person vignette experiment. Participants imagined themselves in a story involving a follower they had to evaluate, then answered scales capturing cognitive self-regulation and their intent to employ abusive supervision. Participants were randomly assigned to a follower engaging in neutral or status-threatening behavior. Results from two-stage path analysis revealed that, in the neutral condition, unidimensional narcissism's impact on abusive supervision intentions was fully mediated by revenge motivations. In the threat condition, narcissism's impact on abusive supervision intentions was partially mediated by revenge motivations and moral disengagement. Threat condition did not significantly strengthen the relationships between narcissism and cognitive self-regulatory responses. A research question tested how tripartite narcissism related to self-regulation variables in the threat condition. Findings suggested self-important narcissism increased revenge motivations and moral disengagement, as well as decreased job performance ratings of the follower. Direct effect estimates revealed grandiose narcissism negatively predicted and self-important narcissism positively predicted abusive supervision intentions. Vulnerable narcissism had no significant impact on any endogenous variables. I conclude that, whether the threat is real or imagined, derogation of others is a key tool narcissists use to defend their status.
  • Building Trustworthy Machine Learning Systems for Security: From Federated Learning to Agentic AI
    Zhang, Chaoyu (Virginia Tech, 2026-08-07)
    Federated Learning (FL) has emerged as a promising paradigm for privacy-preserving collaborative machine learning, enabling distributed devices to jointly train models without sharing raw data. This dissertation addresses critical challenges in building trustworthy and secure machine learning systems across two complementary frontiers: FL systems for network and medical security, and anomaly detection for agentic AI. Together, the five technical chapters form a progression from improving FL's utility and resilience, through applying FL to network security, to defending FL from model inversion attacks in medical settings, and finally to detecting workflow-level anomalies in modern agentic AI systems. Chapter 2 tackles data quality heterogeneity in FL to improve global model utility. Noisy-labeled and imbalanced local data among clients can severely hinder training efficiency and model convergence. We propose a quality and fairness-aware client selection mechanism based on a novel Quality-of-Model (QoM) metric that evaluates client contribution without requiring access to model updates or gradients. Our approach prioritizes clients with high-quality contributions while ensuring diversity through a sortition-inspired randomized selection process, improving convergence speed and reducing communication overhead. Chapter 3 addresses Byzantine resilience to ensure FL trustworthiness from a system security perspective. FL's distributed nature forces the central server to blindly trust local training processes, making it vulnerable to model poisoning and data poisoning attacks from malicious participants. We propose a remote attestation-based approach that regains transparency into client-side training by verifying computation integrity via Trusted Execution Environment (TEE)-generated cryptographic attestation reports, enabling the server to reject malicious updates while preserving performance under non-IID data distributions. Building on these foundations, Chapter 4 demonstrates FL applied to network intrusion detection. We propose a geometric feature learning approach that projects network traffic into a compact, well-structured representation space, combining contrastive feature learning with H-Score optimization to maximize intra-class compactness and inter-class separability. The resulting federated intrusion detection system supports both anomaly detection and precise attack type identification, including zero-day threat exploration through entropy-based uncertainty analysis. Chapter 5 addresses a distinct and critical privacy threat in FL: model inversion attacks (MIAs) that allow a malicious server to reconstruct private training data from shared model updates. We identify that the reconstruction success of all known MIAs is fundamentally bounded by the local batch size relative to the model's first-layer leakage capacity. Building on this observation, we propose Aegis, a defense that synthesizes an auxiliary dataset to push the effective batch size beyond this capacity, collapsing the server's closed-form reconstructions without modifying the FL protocol or perturbing patient data. We demonstrate the approach on medical imaging datasets. Chapter 6 extends our security focus beyond FL to the rapidly growing domain of agentic AI. Modern agentic systems execute complex tasks through long-horizon workflows involving multi-agent coordination and tool invocation, creating a new risk surface where a single injected or erroneous step propagates through downstream dependencies. We present Skynet, a workflow-level anomaly detection framework that models multi-agent execution as directed workflow graphs and learns benign behavior jointly over semantic and structural dimensions. Trained exclusively on benign workflows, Skynet detects both adversarial manipulations and intrinsic execution failures under a single decision rule, naturally extending to zero-day anomalies. Together, these five chapters provide a comprehensive framework for building trustworthy machine learning systems: improving FL utility and trustworthiness, applying FL to network and medical security, and detecting anomalies in emerging agentic AI architectures, advancing the state of the art in machine learning security in adversarial and distributed environments.
  • Exploring Generative AI for Pedagogically Aligned Learning Experiences and Adapted Instructional Practices in Software Engineering Education
    Wang, Tianjia (Virginia Tech, 2026-08-05)
    In the era of generative artificial intelligence(AI), educators and students face both practical opportunities and pressing challenges for software engineering (SE) education. Generative AI powered by large language models is capable of completing complex software development tasks, becoming increasingly integrated into the development process, and is reshaping how students can learn to design, develop, and test software systems. With the advancement of generative AI, it is important to understand how students and educators perceive its role, benefits, and challenges in educational contexts. Also, many existing generative AI tools are adapted from general-purpose models without considering how they align with curriculum goals, cognitive development, or instructional strategies in SE education. In this dissertation, we present our examination of generative AI in SE education from: (1) understanding the perceptions, practices, and expectations of students and instructors regarding generative AI; (2) exploring how intelligent systems powered by generative AI can align with pedagogical goals and support active, collaborative and engaging learning environments; (3) investigating how generative AI could reshape SE knowledge areas to guide curriculum design and assessment. The findings show that generative AI can solve course problem sets with high accuracy, raising concerns about academic integrity, misinformation, and shallow learning; however, instructors remain open to its use when supported by clear policies and redesigned assessments. Students valued generative AI for providing real-time feedback, decomposing complex tasks, and reducing anxiety among those hesitant to ask questions, but identified limitations in emotional connection, nonverbal communication, and human-like interaction. In investigating how generative AI can effectively support student learning while addressing misuse and related issues, the studies on intelligent systems demonstrate that generative AI agents can simulate believable behaviors in learning environments, increasing students' perceived cognitive and social presence. The findings further indicate that generative AI is more effective in supporting student learning and engagement when learning objectives are embedded in system design and AI agents' behaviors are aligned with pedagogical goals. The AI-powered multi-agent system can provide structured, continuous, and interactive scaffolding to help students learn the software development life cycle, significantly improving students' learning gains, increasing task completion rates, and strengthening students' engagement. Finally, the dissertation provides empirical findings that indicate inconsistent topic coverage in undergraduate SE courses. Developers perceived GenAI as most effective for tasks in coding-related knowledge areas. At the same time, participants did not view GenAI as replacing foundational SE knowledge and identified a shift in the importance of the SE knowledge areas. Developers also emphasized the need to include AI-related knowledge and skills in SE education, suggesting that SE curricula should adapt toward preparing students to engineer, validate, and take responsibility for software development produced through human-AI collaboration.
  • Looking for New Particle Physics with Astrophysical Origin
    Gustafson, Robert Andrew (Virginia Tech, 2026-08-05)
    In this thesis, I explore the consequences of introducing new particles into astrophysical environments, and place constraints on these particles using available data. I first consider Heavy Neutral Leptons (a proposed particle which has important implications for neutrinos) in the context of atmospheric interactions, the Sun, and supernovae. I then turn my focus to various models of dark matter, considering the reach of both the terrestrial and astronomical observables. A special focus is given to cases where dark matter clusters around Supermassive Black Holes.
  • Automated Testing for Data-Intensive Scalable Computing
    Humayun, Ahmad (Virginia Tech, 2026-08-05)
    Data-Intensive Scalable Computing (DISC) systems such as Apache Spark, Hadoop, Flink, and Beam have become central to modern data processing. These systems enable developers to write applications as dataflows composed of user-defined functions operating over large and often unstructured datasets. However, testing such applications and the underlying frameworks remains a significant challenge. Conventional testing techniques struggle in this space due to the complexity of dataflow semantics, the scale and heterogeneity of input data, and the need to reason about operator interactions, schema variability, and program logic simultaneously. This thesis presents a set of techniques for extracting and leveraging rich, fine-grained properties that are unique to DISC workloads in order to inform automated input generation for effective testing across various levels of the data-centric software stack. The first thrust of this work focuses on generating inputs that can avoid trivial parsing errors and effectively exercise the deeper logic in DISC applications. By analyzing how code interacts with different parts of the input data, we test the code where it matters most instead of wasteful fuzzing cycles finding parsing issues. The second thrust explores how to create realistic and meaningful test data that reflects the structure and semantics of real-world inputs, while still achieving high coverage and fault detection. The final part of this work shifts focus to DISC frameworks, recognizing that applications are only as reliable as the systems they run on. By generating diverse dataflow programs, we systematically test internal components like optimizers, extending fuzzing to the entire DISC stack.
  • Explainable Machine Learning Framework for Accurate Reference-free Biological Data Deconvolution
    Du, Dongping (Virginia Tech, 2026-08-05)
    Bulk omics data captures mixed molecular signals from multicellular tissue compositions, posing significant challenges in accurate interpretation of biological changes over samples. Reference free deconvolution offers a flexible framework to estimate cell type proportions and specific expressions from bulk data, thus uncovering latent cellular and molecular architecture of tissue ecosystems without relying on predefined references. However, existing deconvolution methods are highly sensitive to several hidden confounders, including asymmetric gene expression, informative missingness, inter-cell-type imbalance, and deviation from identifiability conditions. These issues also propagate across preprocessing and modeling stages, collectively, leading to reduced accuracy of bulk deconvolution and downstream inference. In this dissertation, we present a comprehensive methodological framework with effective workflow for reference-free deconvolution of complex biological data. The proposed workflow integrates four key components spanning preprocessing, missing value imputation, structural correction, and discriminative deconvolution. First, we propose Cosbin, an iterative normalization strategy that identifies consistently expressed genes and removes asymmetrically differentially expressed genes, thereby preserving the geometric structure of the data. Second, we propose mechanism-integrated group-wise pre-imputation, which explicitly models multiple missingness mechanisms and preserves biologically informative missing patterns, particularly for marker-like genes. Third, we introduce iterative equilibration of cell-type expression profiles to correct inter cell-type asymmetry and improve the identifiability and accuracy of proportion estimation. Fourth, we apply and evaluate CAM3.0, an enhanced convex geometry-based unsupervised deconvolution algorithm, to estimate latent molecular archetypes and compositions on diverse omics data types from real biological bulk samples. To further improve deconvolution accuracy and efficiency particularly when constituent cell types are highly mixed or hardly separable, we also propose and develop an effective cosine similarity based discriminative analysis of mixtures method (csDAM), specifically to deconvolute highly mixed bulk expression data. Facilitated by the monotonic relationship between signature gene specificity and cosine similarity rank distribution, csDAM achieves highly accurate estimation of cell type proportions and specific expressions. Simulation and real-data based studies show both improved deconvolution performance and computational efficiency by csDAM compared to most relevant peer methods. In summary, the work in this dissertation addresses several major limitations of existing reference free deconvolution approaches by collectively integrating improved preprocessing, missing value imputation, structural correction, and discriminative modeling into a unified framework. Extensive simulations and real-data applications demonstrate improved performance in terms of stability, accuracy, and interpretability under challenging conditions, including high noise, complex missingness, and strong cellular heterogeneity. This study highlights the critical role of structural considerations in deconvolution and provides a scalable solution for extracting biologically meaningful latent features and archetypes from large-scale bulk omics data.
  • Computational Framework and Deep Learning for Cross-Species Comparison in Plant Transcriptomics and Imaging
    Chau, Tran Ngoc (Virginia Tech, 2026-08-05)
    Cross-species cell type analysis is fundamental to comparative plant biology, enabling the transfer of knowledge from well-studied model species such as Arabidopsis thaliana to agriculturally and medicinally important non-model species. Recent advances in single-cell RNA sequencing (scRNA-seq) have generated large-scale transcriptomic atlases across diverse plant species, yet extracting biological insight from these datasets, and comparing them across species, remains challenging due to data sparsity, dropout noise, limited marker gene information in non-model plants, limited ortholog detection across distant species, and the lack of integrated frameworks across data modalities. This dissertation develops computational frameworks for cross-species cell type mapping in plants, spanning both transcriptomic and imaging data, through a sequence of connected projects. First, we address the lack of reliable tools for identifying co-expressed gene modules and imputing missing values to improve data quality in plant scRNA-seq. We construct a benchmark that evaluates co-expression and imputation methods against a ground truth of promoter-reporter-validated native gene pairs, providing the community with a trusted foundation for downstream analyses such as identifying transcription factor target genes, nuclear pore complex gene associations, and functional gene modules. Second, we introduce the Orthologous Marker Gene (OMG) framework, which maps cell types across 15 plant species by identifying cell-type-specific marker genes that belong to conserved orthogroups. Rather than relying on one-to-one gene-level orthology, which is confounded by sequence divergence across distantly related species, OMG aggregates markers at the orthogroup level, enabling robust cross-species cell type correspondence. Then, we extend this framework with PLM-OMG, a protein language model-based approach for scalable orthogroup classification, enabling incremental inclusion of new species without recomputing existing orthogroups, substantially reducing computational overhead and improving cross-species cell type mapping. Finally, we extend cross-species cell type analysis to imaging by developing a hybrid pipeline for confocal root images across different species and developmental stages. The pipeline integrates Cellpose-SAM segmentation with a classifier that combines morpho-topological features and fine-tuned DINOv2 vision transformer embeddings, trained using gradient-boosted models and iterative refinement. Ambiguous cells are resolved using a vision-language model, demonstrating scalable automated cell type annotation across diverse plant species. Together, these contributions provide reusable computational frameworks for cross-species cell type mapping in plants, spanning transcriptomic and imaging modalities, and advancing comparative plant genomics at scale.
  • From Protest to Policy: An AI-Assisted Meta-Analytic Journey Through LGBTQ+ Change
    Cornett, Kelsi (Virginia Tech, 2026-07-31)
    This dissertation examines how activism, public opinion, and legal change interact in the advancement of LGBTQ+ rights. A human-in-the-loop, AI-assisted meta-analytic workflow employed ASReview LAB, Elicit AI, and GPT to support screening, data extraction, and validation. The final synthesis included 28 effects represented in 25 independent study matrices with a summed analytic sample size of N=1,613,593. Random-effects meta-analytic structural equation modeling (MASEM) was used to pool associations among the three constructs and compare competing conceptual models of sociopolitical change. Activism and public opinion were each positively associated with policy or legal change, whereas the direct association between activism and public opinion was small and nonsignificant. The strongest-fitting model was the policy-responsive mobilization model (Model 9; CFI= 1.00, TLI= 1.278, AIC= -1.98, BIC=-14.27). In this model, public opinion predicts policy or legal change (b = 0.186), which subsequently predicts activism (b = 0.187). However, the indirect effect was not statistically significant, so the findings do not confirm mediation or causality. The public opinion-policy relationship was the most stable across sensitivity and influence analyses. Nevertheless, the evidence base was highly heterogeneous, temporally constrained to the 21st century, concentrated in the Global North, and unevenly distributed across relationships. Overall, the findings suggest that legal reform may function as both an outcome of social movements and as a source of legitimacy and opportunity for renewed activism.
  • Explaining and Steering Embedding Projections for Visual Analytics
    Liu, Wei (Virginia Tech, 2026-07-29)
    Low-dimensional embedding projections are widely used in visual analytics, particularly for exploring large document collections. By arranging documents as points in a two-dimensional space, these projections help analysts identify clusters, outliers, separations, and relationships among documents. However, projection layouts are often difficult to interpret and control: users can observe where documents are positioned, but may not understand why spatial patterns appear or how to reshape the projection when the resulting layout does not align with their analytic goals. This dissertation frames embedding projections as interactive semantic workspaces and develops methods for explaining and steering them in visual analytics. First, it introduces gradient-based explanations that connect textual features to document positions in projection layouts, revealing how words influence spatial placement. Second, it presents context-aware natural-language explanations that combine document semantics with layout-derived spatial context to help users interpret documents, regions, and spatial patterns. Third, it moves from explanation to steering by introducing an LLM-augmented semantic steering approach, in which analysts express semantic intent through example groupings and reshape projections without retraining the underlying models. Finally, it develops a scalable prototype-based steering method that shifts LLM reasoning from individual items to group-level abstraction, making semantic steering practical for large embedding collections. Through quantitative evaluations, usage scenarios, case studies, and a user study, this dissertation demonstrates that embedding projections can be made more interpretable, controllable, and aligned with analytic goals. These contributions advance projection-based visual analytics toward interactive semantic workspaces that analysts can inspect, understand, and reshape.
  • Towards Formally Verified Invariant Properties of Control Code Using Robustness Analysis Results
    Khalife, Elias (Virginia Tech, 2026-07-29)
    Safety-critical aerospace systems rely on digital feedback controllers designed to satisfy stability, performance, and safety requirements despite uncertainties and disturbances. These guarantees are typically established through robust control analysis performed at the model level, under the assumption of exact, infinite-precision arithmetic. However, such guarantees do not automatically extend to the executable code, as floating-point computation and roundoff errors can compromise model-level properties. This dissertation addresses this gap by establishing and connecting robust control analysis results for discrete-time uncertain systems with the deductive formal verification of the corresponding executable code. Robust control analysis results, however, can be conservative. One way to assess this conservatism is to examine how a system behaves under worst-case conditions. To this end, this dissertation develops methods for constructing input signals that drive asymptotically stable, discrete-time linear time-varying systems toward their worst-case behavior, providing a means of evaluating the tightness of established performance results. Beyond this assessment, the dissertation advances the underlying analysis frameworks, including methods for computing invariant and bounding ellipsoids for systems with uncertainties characterized using pointwise integral quadratic constraints. It also establishes stability and performance guarantees for switched linear control laws. To further reduce conservatism in reachability analysis, sum-of-squares programming is used to construct verified quadratic characterizations for smooth nonlinearities and activation functions in feedforward neural networks. Ultimately, these results yield ellipsoidal and quadratic-constraint certificates that establish invariant properties at the model level. On the verification side, this dissertation establishes a deductive workflow that translates these mathematical certificates into code-level contracts, formally verifying the corresponding executable C code while explicitly accounting for floating-point errors. This workflow is then extended from local, single-component properties to systemic, closed-loop properties. By encoding an augmented-system abstraction using ghost code, the approach enables the formal verification of these systemic properties in both the real and float models, without modifying the executable controller code. Collectively, this dissertation establishes a certificate-based pathway from robust control analysis to the formal verification of control programs, while also developing the tools needed to assess and reduce the conservatism of the underlying analysis.
  • Topographic and depth-to-bedrock controls on non-perennial headwater streamflow
    Morgan, John Cole (Virginia Tech, 2026-07-28)
    Non-perennial headwater streams dominate river networks in terms of total flowing length and contribute substantially to downstream water conditions. The flowing extent of these streams varies over space and time, driven by several interacting controls that are themselves difficult to predict. Two of these major controls are topography and subsurface characteristics. The goal of this dissertation was to understand how topography and the distribution of subsurface storage, represented by depth-to-bedrock (DTB), interact to control patterns of non-perennial streamflow at the Hubbard Brook Experimental Forest in the White Mountains of New Hampshire. To do this, I conducted a field campaign to collect flowing-state observations for three non-perennial streams using a network of distributed flow presence and absence sensors. I paired these observations with a characterization of the depth-to-bedrock across one study watershed, combining passive seismic sensing with direct observations from ground surveys and soil pits. These datasets were integrated with a suite of topographic analyses and modeling methods, to examine how topography and depth-to-bedrock interact to explain wetting and drying patterns and flow persistence of a stream network. This dissertation had three research objectives: 1) to determine the order of wetting and drying in non-perennial stream networks during events, and if it can reveal catchment subsurface structure 2) to test whether characterization of depth-to-bedrock throughout a watershed explains flow persistence in a stream network, and 3) evaluating whether a process-based model with additional subsurface information can reproduce observed surface flow dynamics in space and time. Across these methods, both topography and subsurface properties explain processes that drive temporary flow in non-perennial streams. Non-perennial headwater streams do not activate in a fixed order of wetting and drying from event to event. There is no strong relationship between depth-to-bedrock and flow persistence, either locally or across hillslopes that drain to channel reaches. The process-based model was the most transferable approach, producing accurate predictions of discharge at the watershed outlet, but it could not resolve the finer-scale patterns of wetting and drying throughout the network. Explaining the controls of wetting and drying in non-perennial streams remains challenging, but this work validates past findings that topography-based frameworks are easy to implement and effective, and shows that, at least in these headwater networks, site-specific knowledge and thorough characterization of the subsurface does not improve the ability to predict these small, complex systems.
  • Policy, Practice, and Possibility: A Collective Case Study Exploring Agency Among Veteran School-Based Agriculture Educators in Virginia
    Taylor, Demikia Surgeon (Virginia Tech, 2026-07-22)
    Educational policy is rarely shaped by those who carry the role of implementation. Due to existing and impending educational policy structures, CTE teachers, specifically in agriculture education, must navigate these systems to maintain their school-based agriculture education (SBAE) programs. This study examined how veteran SBAE teachers in Virginia perceive and enact agency as they navigate local and state-level educational policy structures. Grounded in the Teacher Agency Model (TAM) (Priestley et al., 2015) and analyzed through the lens of Critical Policy Analysis (CPA) framework (Diem et al., 2014; Diem and Young, 2018), this study investigated existing educational policies targeting SBAE and career and technical education (CTE) programs to ensure equitable access is applied, and to determine how teachers use their institutional knowledge to showcase their agentic practice through advocacy. Using a collective case study design with an ethnographic approach, data was collected through policy document analysis, participant observations, and semi-structured interviews. Six themes emerged from the data analysis, which were organized into TAM's three dimensions: Iterational, Practical-Evaluative, and Projective. The themes that emerged included: 1) Valuing Institutional Knowledge and Experience, 2) Exclusion from Policy Decision-Making, 3) Leadership Knowledge Gaps and Resource Restrictions, 4) Balancing Workloads and Unrealistic Expectations, 5) Leveraging Visibility for Programmatic Change, and 6) Supportive Relationships Fostering Agentic Practice. Findings revealed that SBAE teachers were empowered advocates and agents of change for their programs, using their professional knowledge and strategic relationship-building to address systemic exclusionary practices that hinder the development and sustainability of their SBAE programs.
  • Designing Interactive Systems to Facilitate Exploration and Participation in Live Coding
    Manesh, Daniel Madjid (Virginia Tech, 2026-07-17)
    In computing education, live coding refers to a lecture technique in which an instructor writes code in front of students as a part of the lecture, speaking aloud throughout to provide commentary and explain their thought process. In the performing arts, live coding is a practice in which a performer continuously writes, edits, and runs code live in front of an audience to create an audiovisual performance. In this dissertation, I explore how we can design systems to support and augment the practice of live coding in both domains by addressing two key challenges: exploration and participation. In the performing arts domain, I present SHARP, a lightweight, block-level version control system designed for the live coding music language Tidal Cycles. A user study revealed that SHARP's version trees made it easier to understand the progression of code and enabled live coding musicians to explore new combinations of sounds and to execute musical forms on the fly. In the computing education domain, I conducted 22 interviews with CS instructors who use live coding and discovered that (1) instructors had conflicting opinions on whether students should type along with them during live coding; and (2) instructors valued that live coding offered many touchpoints for student participation, but wished that students participated more. Regarding (1), I designed and developed NOTES, a note-taking system for live coding that allows students to take snapshots of the instructor's code, enabling students to focus on writing their own notes rather than just copying the instructor's code. Regarding (2), I designed and developed AD LIB, a system that enables instructors to quickly create class exercises on the fly while live coding, encouraging student participation and enabling controlled student exploration through coding activities. Finally, I discuss how each system is united by a common theme of version control, and I explore avenues for future work designing interactive systems for understanding how code evolves over time.
  • Supervised Variational Autoencoders for Structural Learning and Statistical Inference with Heterogeneous Data
    Lee, Jaeyoung (Virginia Tech, 2026-07-14)
    Large-scale datasets, such as images, often exhibit heterogeneous structures caused by diverse subpopulations or complex experimental designs. Extracting meaningful low-dimensional representations from such data, while accounting for heterogeneity and enabling statistical inference, remains a significant challenge. This dissertation addresses two related challenges within variational autoencoder (VAE) frameworks. The first project introduces the Generalized Variational Autoencoder (GVAE), a unified deep generative model that integrates dimension reduction, structural learning, and adaptive prediction into a single framework. GVAE constructs a composite latent space consisting of a Gaussian component for feature extraction and a stick-breaking process component for capturing latent subpopulation structure. A mixture-of-experts predictor linked to this composite latent space produces predictions that adapt to the subgroups. Simulation studies and an application to brain tumor MRI data demonstrate that GVAE achieves improved predictive accuracy over competing models while effectively recovering the underlying latent heterogeneous structure. The conditional generative component further reveals how images vary with the response, providing an additional tool for understanding complex data. The second project extends the VAE framework to association testing by proposing a two-step M-estimation approach in which VAE encoder parameters serve as first-step nuisance estimators for testing the association between high-dimensional inputs and a scalar response variable. We establish two Wilks-type asymptotic results depending on how the first-step nuisance parameters are obtained. When the nuisance parameters are trained with an unsupervised VAE, the classical Wilks' theorem holds and the two-step likelihood ratio statistic converges to a chi-squared distribution under the null. When the parameters are trained with a supervised VAE, the classical approximation may fail and we conjecture that a scaled likelihood ratio statistic still follows an approximate chi-squared distribution with adjusted degrees of freedom. Through simulation studies, we investigate the asymptotic behavior of the proposed test statistic and evaluate the type I error and power.
  • Statistical Methods for Artificial Intelligence Reliability with Applications in Autonomous Vehicles
    Zheng, Simin (Virginia Tech, 2026-07-14)
    Recurrent event data provide an important source of information for assessing the reliability of artificial intelligence (AI) systems. As AI technologies continue to be adopted in safety-critical applications, there is a growing need for statistical methodologies that can effectively model, test, and assure their reliability. This dissertation develops statistical methods for AI reliability analysis using recurrent event data, with a particular focus on autonomous vehicles (AVs). The proposed research addresses reliability evaluation across multiple stages of the AI system life cycle, including reliability modeling, accelerated testing, and reliability assurance. First, a statistical recurrent-event modeling framework with multivariate random effects is developed to analyze multiple types of recurrent events simultaneously. The proposed approach accounts for dependence among recurrent-event processes while accommodating unit-level heterogeneity. Second, a screening-based accelerated life testing methodology incorporating regularization techniques is proposed for AI systems with recurrent-event responses. The proposed framework utilizes a regularized semiparametric recurrent-event model with random effects to identify influential driving factors, interaction effects, and potential nonlinear relationships, providing an efficient strategy for reliability evaluation during the development stage. Third, a statistical framework for reliability assurance test planning based on recurrent-event processes is developed to support deployment decisions. The proposed methods establish testing requirements and decision criteria for demonstrating reliability before field operation. The proposed methodologies are evaluated through simulation studies and illustrated using publicly available AV data from the California Department of Motor Vehicles testing program, as well as recurrent-event data generated from physics-based simulation environments such as CARLA. Overall, this dissertation provides a comprehensive statistical framework for AI reliability analysis that supports the safe development, evaluation, and deployment of AVs and other AI-enabled systems.