Doctoral Dissertations

Permanent URI for this collection

Browse

Recent Submissions

Now showing 1 - 20 of 18527
  • From Protest to Policy: An AI-Assisted Meta-Analytic Journey Through LGBTQ+ Change
    Cornett, Kelsi (Virginia Tech, 2026-07-31)
    This dissertation examines how activism, public opinion, and legal change interact in the advancement of LGBTQ+ rights. A human-in-the-loop, AI-assisted meta-analytic workflow employed ASReview LAB, Elicit AI, and GPT to support screening, data extraction, and validation. The final synthesis included 28 effects represented in 25 independent study matrices with a summed analytic sample size of N=1,613,593. Random-effects meta-analytic structural equation modeling (MASEM) was used to pool associations among the three constructs and compare competing conceptual models of sociopolitical change. Activism and public opinion were each positively associated with policy or legal change, whereas the direct association between activism and public opinion was small and nonsignificant. The strongest-fitting model was the policy-responsive mobilization model (Model 9; CFI= 1.00, TLI= 1.278, AIC= -1.98, BIC=-14.27). In this model, public opinion predicts policy or legal change (b = 0.186), which subsequently predicts activism (b = 0.187). However, the indirect effect was not statistically significant, so the findings do not confirm mediation or causality. The public opinion-policy relationship was the most stable across sensitivity and influence analyses. Nevertheless, the evidence base was highly heterogeneous, temporally constrained to the 21st century, concentrated in the Global North, and unevenly distributed across relationships. Overall, the findings suggest that legal reform may function as both an outcome of social movements and as a source of legitimacy and opportunity for renewed activism.
  • Explaining and Steering Embedding Projections for Visual Analytics
    Liu, Wei (Virginia Tech, 2026-07-29)
    Low-dimensional embedding projections are widely used in visual analytics, particularly for exploring large document collections. By arranging documents as points in a two-dimensional space, these projections help analysts identify clusters, outliers, separations, and relationships among documents. However, projection layouts are often difficult to interpret and control: users can observe where documents are positioned, but may not understand why spatial patterns appear or how to reshape the projection when the resulting layout does not align with their analytic goals. This dissertation frames embedding projections as interactive semantic workspaces and develops methods for explaining and steering them in visual analytics. First, it introduces gradient-based explanations that connect textual features to document positions in projection layouts, revealing how words influence spatial placement. Second, it presents context-aware natural-language explanations that combine document semantics with layout-derived spatial context to help users interpret documents, regions, and spatial patterns. Third, it moves from explanation to steering by introducing an LLM-augmented semantic steering approach, in which analysts express semantic intent through example groupings and reshape projections without retraining the underlying models. Finally, it develops a scalable prototype-based steering method that shifts LLM reasoning from individual items to group-level abstraction, making semantic steering practical for large embedding collections. Through quantitative evaluations, usage scenarios, case studies, and a user study, this dissertation demonstrates that embedding projections can be made more interpretable, controllable, and aligned with analytic goals. These contributions advance projection-based visual analytics toward interactive semantic workspaces that analysts can inspect, understand, and reshape.
  • Towards Formally Verified Invariant Properties of Control Code Using Robustness Analysis Results
    Khalife, Elias (Virginia Tech, 2026-07-29)
    Safety-critical aerospace systems rely on digital feedback controllers designed to satisfy stability, performance, and safety requirements despite uncertainties and disturbances. These guarantees are typically established through robust control analysis performed at the model level, under the assumption of exact, infinite-precision arithmetic. However, such guarantees do not automatically extend to the executable code, as floating-point computation and roundoff errors can compromise model-level properties. This dissertation addresses this gap by establishing and connecting robust control analysis results for discrete-time uncertain systems with the deductive formal verification of the corresponding executable code. Robust control analysis results, however, can be conservative. One way to assess this conservatism is to examine how a system behaves under worst-case conditions. To this end, this dissertation develops methods for constructing input signals that drive asymptotically stable, discrete-time linear time-varying systems toward their worst-case behavior, providing a means of evaluating the tightness of established performance results. Beyond this assessment, the dissertation advances the underlying analysis frameworks, including methods for computing invariant and bounding ellipsoids for systems with uncertainties characterized using pointwise integral quadratic constraints. It also establishes stability and performance guarantees for switched linear control laws. To further reduce conservatism in reachability analysis, sum-of-squares programming is used to construct verified quadratic characterizations for smooth nonlinearities and activation functions in feedforward neural networks. Ultimately, these results yield ellipsoidal and quadratic-constraint certificates that establish invariant properties at the model level. On the verification side, this dissertation establishes a deductive workflow that translates these mathematical certificates into code-level contracts, formally verifying the corresponding executable C code while explicitly accounting for floating-point errors. This workflow is then extended from local, single-component properties to systemic, closed-loop properties. By encoding an augmented-system abstraction using ghost code, the approach enables the formal verification of these systemic properties in both the real and float models, without modifying the executable controller code. Collectively, this dissertation establishes a certificate-based pathway from robust control analysis to the formal verification of control programs, while also developing the tools needed to assess and reduce the conservatism of the underlying analysis.
  • Topographic and depth-to-bedrock controls on non-perennial headwater streamflow
    Morgan, John Cole (Virginia Tech, 2026-07-28)
    Non-perennial headwater streams dominate river networks in terms of total flowing length and contribute substantially to downstream water conditions. The flowing extent of these streams varies over space and time, driven by several interacting controls that are themselves difficult to predict. Two of these major controls are topography and subsurface characteristics. The goal of this dissertation was to understand how topography and the distribution of subsurface storage, represented by depth-to-bedrock (DTB), interact to control patterns of non-perennial streamflow at the Hubbard Brook Experimental Forest in the White Mountains of New Hampshire. To do this, I conducted a field campaign to collect flowing-state observations for three non-perennial streams using a network of distributed flow presence and absence sensors. I paired these observations with a characterization of the depth-to-bedrock across one study watershed, combining passive seismic sensing with direct observations from ground surveys and soil pits. These datasets were integrated with a suite of topographic analyses and modeling methods, to examine how topography and depth-to-bedrock interact to explain wetting and drying patterns and flow persistence of a stream network. This dissertation had three research objectives: 1) to determine the order of wetting and drying in non-perennial stream networks during events, and if it can reveal catchment subsurface structure 2) to test whether characterization of depth-to-bedrock throughout a watershed explains flow persistence in a stream network, and 3) evaluating whether a process-based model with additional subsurface information can reproduce observed surface flow dynamics in space and time. Across these methods, both topography and subsurface properties explain processes that drive temporary flow in non-perennial streams. Non-perennial headwater streams do not activate in a fixed order of wetting and drying from event to event. There is no strong relationship between depth-to-bedrock and flow persistence, either locally or across hillslopes that drain to channel reaches. The process-based model was the most transferable approach, producing accurate predictions of discharge at the watershed outlet, but it could not resolve the finer-scale patterns of wetting and drying throughout the network. Explaining the controls of wetting and drying in non-perennial streams remains challenging, but this work validates past findings that topography-based frameworks are easy to implement and effective, and shows that, at least in these headwater networks, site-specific knowledge and thorough characterization of the subsurface does not improve the ability to predict these small, complex systems.
  • Policy, Practice, and Possibility: A Collective Case Study Exploring Agency Among Veteran School-Based Agriculture Educators in Virginia
    Taylor, Demikia Surgeon (Virginia Tech, 2026-07-22)
    Educational policy is rarely shaped by those who carry the role of implementation. Due to existing and impending educational policy structures, CTE teachers, specifically in agriculture education, must navigate these systems to maintain their school-based agriculture education (SBAE) programs. This study examined how veteran SBAE teachers in Virginia perceive and enact agency as they navigate local and state-level educational policy structures. Grounded in the Teacher Agency Model (TAM) (Priestley et al., 2015) and analyzed through the lens of Critical Policy Analysis (CPA) framework (Diem et al., 2014; Diem and Young, 2018), this study investigated existing educational policies targeting SBAE and career and technical education (CTE) programs to ensure equitable access is applied, and to determine how teachers use their institutional knowledge to showcase their agentic practice through advocacy. Using a collective case study design with an ethnographic approach, data was collected through policy document analysis, participant observations, and semi-structured interviews. Six themes emerged from the data analysis, which were organized into TAM's three dimensions: Iterational, Practical-Evaluative, and Projective. The themes that emerged included: 1) Valuing Institutional Knowledge and Experience, 2) Exclusion from Policy Decision-Making, 3) Leadership Knowledge Gaps and Resource Restrictions, 4) Balancing Workloads and Unrealistic Expectations, 5) Leveraging Visibility for Programmatic Change, and 6) Supportive Relationships Fostering Agentic Practice. Findings revealed that SBAE teachers were empowered advocates and agents of change for their programs, using their professional knowledge and strategic relationship-building to address systemic exclusionary practices that hinder the development and sustainability of their SBAE programs.
  • Designing Interactive Systems to Facilitate Exploration and Participation in Live Coding
    Manesh, Daniel Madjid (Virginia Tech, 2026-07-17)
    In computing education, live coding refers to a lecture technique in which an instructor writes code in front of students as a part of the lecture, speaking aloud throughout to provide commentary and explain their thought process. In the performing arts, live coding is a practice in which a performer continuously writes, edits, and runs code live in front of an audience to create an audiovisual performance. In this dissertation, I explore how we can design systems to support and augment the practice of live coding in both domains by addressing two key challenges: exploration and participation. In the performing arts domain, I present SHARP, a lightweight, block-level version control system designed for the live coding music language Tidal Cycles. A user study revealed that SHARP's version trees made it easier to understand the progression of code and enabled live coding musicians to explore new combinations of sounds and to execute musical forms on the fly. In the computing education domain, I conducted 22 interviews with CS instructors who use live coding and discovered that (1) instructors had conflicting opinions on whether students should type along with them during live coding; and (2) instructors valued that live coding offered many touchpoints for student participation, but wished that students participated more. Regarding (1), I designed and developed NOTES, a note-taking system for live coding that allows students to take snapshots of the instructor's code, enabling students to focus on writing their own notes rather than just copying the instructor's code. Regarding (2), I designed and developed AD LIB, a system that enables instructors to quickly create class exercises on the fly while live coding, encouraging student participation and enabling controlled student exploration through coding activities. Finally, I discuss how each system is united by a common theme of version control, and I explore avenues for future work designing interactive systems for understanding how code evolves over time.
  • Supervised Variational Autoencoders for Structural Learning and Statistical Inference with Heterogeneous Data
    Lee, Jaeyoung (Virginia Tech, 2026-07-14)
    Large-scale datasets, such as images, often exhibit heterogeneous structures caused by diverse subpopulations or complex experimental designs. Extracting meaningful low-dimensional representations from such data, while accounting for heterogeneity and enabling statistical inference, remains a significant challenge. This dissertation addresses two related challenges within variational autoencoder (VAE) frameworks. The first project introduces the Generalized Variational Autoencoder (GVAE), a unified deep generative model that integrates dimension reduction, structural learning, and adaptive prediction into a single framework. GVAE constructs a composite latent space consisting of a Gaussian component for feature extraction and a stick-breaking process component for capturing latent subpopulation structure. A mixture-of-experts predictor linked to this composite latent space produces predictions that adapt to the subgroups. Simulation studies and an application to brain tumor MRI data demonstrate that GVAE achieves improved predictive accuracy over competing models while effectively recovering the underlying latent heterogeneous structure. The conditional generative component further reveals how images vary with the response, providing an additional tool for understanding complex data. The second project extends the VAE framework to association testing by proposing a two-step M-estimation approach in which VAE encoder parameters serve as first-step nuisance estimators for testing the association between high-dimensional inputs and a scalar response variable. We establish two Wilks-type asymptotic results depending on how the first-step nuisance parameters are obtained. When the nuisance parameters are trained with an unsupervised VAE, the classical Wilks' theorem holds and the two-step likelihood ratio statistic converges to a chi-squared distribution under the null. When the parameters are trained with a supervised VAE, the classical approximation may fail and we conjecture that a scaled likelihood ratio statistic still follows an approximate chi-squared distribution with adjusted degrees of freedom. Through simulation studies, we investigate the asymptotic behavior of the proposed test statistic and evaluate the type I error and power.
  • Statistical Methods for Artificial Intelligence Reliability with Applications in Autonomous Vehicles
    Zheng, Simin (Virginia Tech, 2026-07-14)
    Recurrent event data provide an important source of information for assessing the reliability of artificial intelligence (AI) systems. As AI technologies continue to be adopted in safety-critical applications, there is a growing need for statistical methodologies that can effectively model, test, and assure their reliability. This dissertation develops statistical methods for AI reliability analysis using recurrent event data, with a particular focus on autonomous vehicles (AVs). The proposed research addresses reliability evaluation across multiple stages of the AI system life cycle, including reliability modeling, accelerated testing, and reliability assurance. First, a statistical recurrent-event modeling framework with multivariate random effects is developed to analyze multiple types of recurrent events simultaneously. The proposed approach accounts for dependence among recurrent-event processes while accommodating unit-level heterogeneity. Second, a screening-based accelerated life testing methodology incorporating regularization techniques is proposed for AI systems with recurrent-event responses. The proposed framework utilizes a regularized semiparametric recurrent-event model with random effects to identify influential driving factors, interaction effects, and potential nonlinear relationships, providing an efficient strategy for reliability evaluation during the development stage. Third, a statistical framework for reliability assurance test planning based on recurrent-event processes is developed to support deployment decisions. The proposed methods establish testing requirements and decision criteria for demonstrating reliability before field operation. The proposed methodologies are evaluated through simulation studies and illustrated using publicly available AV data from the California Department of Motor Vehicles testing program, as well as recurrent-event data generated from physics-based simulation environments such as CARLA. Overall, this dissertation provides a comprehensive statistical framework for AI reliability analysis that supports the safe development, evaluation, and deployment of AVs and other AI-enabled systems.
  • Defending the Dialog: Characterizing and Mitigating Harmful Fine-Tuning Attacks in Conversational AI
    Cheruvu, Aravind (Virginia Tech, 2026-07-14)
    Generative AI (GenAI), especially Large Language Models (LLMs), has transformed conversational AI, enabling chatbots that converse fluently across a wide range of topics and perform diverse downstream tasks. There is growing demand to customize chatbots for specific downstream applications, and LLMs are routinely fine-tuned on application or task-specific conversational datasets. This customization paradigm, however, introduces a new attack surface. We refer to this emerging class of attacks as harmful fine-tuning attacks, in which an adversary poisons the training dataset with harmful or unsafe samples, causing the resulting chatbot to produce unsafe responses. Such attacks can cause real harm, particularly for vulnerable populations, including minorities and those with physical or mental health challenges. This dissertation systematically characterizes and defends against harmful fine-tuning attacks in conversational AI. On the attack characterization side, we focus on toxicity injection attacks, a subclass of harmful fine-tuning attacks in which an adversary injects toxic language into the training dataset. We design malicious agents that leverage GenAI to automatically inject toxic conversations into a dialog-based learning (DBL) pipeline. We then study how such attacks can poison a deployed chatbot. We further investigate the resilience of existing defenses and find that while they are effective in non-adaptive settings, they become vulnerable to adaptive adversarial attacks. Building upon these findings, on the defense side, we propose Optimus, a novel defense framework that mitigates toxicity while preserving conversational utility and seamlessly integrates into existing fine-tuning pipelines. Optimus uses safety-aligned LLMs to identify toxic samples and combines synthetic healing data with direct preference optimization (DPO) to steer the chatbot toward safe responses. Optimus remains effective even when the toxicity classifier is biased or imperfect. Optimus also remains resilient against adaptive adversarial attacks that target individual stages of the defense pipeline. Modern general-purpose chatbots, however, face a much broader range of safety harms beyond toxicity. Building upon Optimus, we propose SafetyOpus, a generalized defense framework that extends its design to a broad spectrum of safety harms spanning hate speech, drug abuse, self-harm, violence, and child abuse. SafetyOpus uses purpose-built LLM-based guardrail filters to identify harmful context-response pairs. SafetyOpus's framework restores the chatbot's safety and utility to levels comparable to those of the original base model, while also mitigating the alignment drift that arises during fine-tuning. Notably, this performance holds even when the underlying guardrail filter is imperfect, demonstrating that SafetyOpus generalizes beyond the filter's detection horizon. This dissertation takes an important step toward making chatbot customization safer by offering deployable defenses and valuable insights for future research toward developing safer and more reliable conversational AI.
  • Understanding Digital Misinformation Across Platforms and Modalities
    Hussein, Eslam Ali Hassan (Virginia Tech, 2026-07-14)
    This dissertation investigates how misinformation spreads across digital platforms and what factors amplify its reach. It examines the role of opaque algorithmic systems—such as those used by YouTube and Amazon—in shaping user exposure to false or misleading content. It also explores how narrative techniques used in social media posts influence public engagement with misinformation, and how manipulated videos, especially those deceptively edited, can be detected using advanced machine learning models. The research is based on four complementary projects that together provide a multi-layered analysis: two algorithmic audits of recommendation systems, a large-scale study of narrative styles in health-related tweets, and a multimodal framework for detecting deceptive video content. Collectively, these studies contribute new datasets, methods, and insights toward understanding and combating misinformation in modern information ecosystems.
  • Pollinators in Virginia Cucurbits
    Walls, Courtney May (Virginia Tech, 2026-07-10)
    Cucurbita pepo plants (pumpkins and squash) are predominately monoecious, requiring pollinators for effective fruit set due to the separation of staminate and pistillate flowers. Consequently, growers often introduce managed pollinators, such as honey bees (Apis mellifera L.) and bumble bees (Bombus spp.), to enhance crop yields. This creates a complex dynamic of interactions between plants, managed pollinators, and native pollinators, particularly a native specialist bee, the squash bee (Xenoglossa pruinosa Say). This work integrates observational methods, analysis of bee microbiome compositions following exposure to C. pepo pollen, and characterization of viral dynamics within the C. pepo agroecosystem to address knowledge gaps regarding constraints on pollinator health and performance in Virginia cucurbit production. Pollinator observations were first conducted to develop a standardized method for quantifying pollinator communities at the field level, enabling growers and experts cooperatively to identify key pollinators and implement management strategies for pollinator protection. This goal was addressed in two chapters. The first demonstrated that visual observations differed from vacuum sampling and bowl trapping in their ability to detect the three primary pollinators in Virginia cucurbits: A. mellifera, Bombus spp., and X. pruinosa. The second chapter evaluated camera trapping relative to visual observations as a tool to facilitate grower and expert collaboration in pollinator identification. Results indicated no significant differences in morphotaxa groups observations of honey bees and bumble bees; however, there were differences in squash bees between methods of observation in the month of August. Subsequent work examined gut bacterial microbiomes of the three major bee types and assessed constraints encountered by social bees (Apis mellifera and Bombus impatiens Cresson) when exposed to C. pepo pollen. We observed a significant decline in the Firmicutes phylum of bacteria, in particular Lactobacillus spp., following consumption of C. pepo pollen in social bees. This suggests C. pepo may indirectly affect overall bee health by altering gut microbial composition and potentially reducing the bee's ability to utilize diverse food sources. Finally, we examined viral communities, including both plant and bee viruses, across cucurbit fields and associated pollinators. Results suggest that all major pollinator types may contribute to horizontal transmission of cucurbit plant viruses, and we identified potential hotspots for bee virus spillover within the C. pepo agroecosystem. Additionally, we report the detection of bee viruses within the cleptoparasite of squash bees and propose potential spillover routes to this parasite. Collectively, these findings highlight the need for further research to assess the ecological risks and benefits of introducing managed bees to Virginia C. pepo systems, for both pollinator and plant health.
  • Investigation of Compact Auxiliary Components for Medium-Voltage Converters with Enhanced Insulation and Wide-Bandgap Device Utilization
    Yan, Ning (Virginia Tech, 2026-07-10)
    Medium-voltage (MV) power electronic systems demand compact, reliable auxiliary and protection hardware that must simultaneously satisfy high insulation voltage, low common-mode (CM) coupling capacitance, and stable operation under fast voltage transients driven by wide-bandgap (WBG) switching devices. This dissertation addresses these competing requirements through three interconnected research directions targeting magnetic isolation, capacitive isolation, and high-voltage solid-state switching for compact MV auxiliary power supplies and switching functions. The first contribution develops a current-transformer-based auxiliary power supply (APS) for distributed MV gate drivers and sensors. An insulation-oriented design methodology is established that co-optimizes transformer geometry, winding placement, conductor structure, and dielectric spacing based on electric-field (E-field) distribution and CM coupling paths. The hardware prototype delivers 20 W output power with 0.83 pF CM capacitance and partial-discharge-free (PD-free) operation up to 14 kV, demonstrating that magnetic isolation can satisfy strict low-coupling requirements when insulation and magnetic performance are co-designed. The second contribution investigates capacitive isolation as a compact, planar alternative for high-frequency MV auxiliary power transfer. A novel multi-dielectric terminal E-field control technique is proposed for PCB-based isolation capacitors to address the dominant insulation failure mode: terminal-edge E-field concentration. By separating the capacitance-forming region from the high-field terminal region and applying material-specific insulation design to each, the proposed extended and encapsulated terminal structure improves PD-free voltage from approximately 4.1 kV to above 38.9 kV—roughly a ten-times improvement in insulation capability. A comparative analysis with the transformer-based design shows that the capacitive structure becomes more volume-efficient beyond approximately 11 kV, with the key tradeoff being higher inherent CM coupling capacitance. Building on this capacitor, a MHz-frequency capacitive-isolated resonant SEPIC converter is developed and experimentally verified, achieving wide voltage-conversion operation, regulated output, over 100 W output power, and a peak efficiency of approximately 87.7%, demonstrating feasibility for practical compact MV auxiliary power delivery. The third contribution presents a SiC MOSFET-based super-cascode switch for MV power switching, capacitor discharge, dc-link chopping, and protection functions. A systematic balancing-network design method is established by separating and analyzing the distinct roles of capacitance offset (ΔC), absolute balancing capacitance (Cₙ), and drain-source snubber capacitance (Cd,ₙ) on turn-on synchronization, turn-off voltage distribution, oscillation suppression, switching speed, device-voltage utilization, and operating-voltage range. A complete turn-off voltage-distribution model is developed to predict final device voltages without assuming ideal sharing, enabling the design to be extended beyond turn-on-only discharge applications. Experimental validation across reduced-device tests, five-device double-pulse tests, and a ten-device hardware prototype demonstrates PD-free operation to 28.1 kV, high-speed switching up to 14 kV at approximately 90.3% device-voltage utilization, and controlled capacitor discharge and dc-link chopping up to a 22 kV dc link using stacked 3.3 kV SiC MOSFETs. Together, these three research directions provide design methods, analytical frameworks, and hardware demonstrations for compact, high-density MV auxiliary and protection hardware. The results quantify and clarify the fundamental tradeoffs among insulation voltage, CM coupling, capacitance density, switching speed, and device-voltage utilization, offering practical design guidance for next-generation medium-voltage power electronic systems.
  • Workload-Aware I/O Optimization in Distributed Object Storage Systems
    Biswas, Debasmita (Virginia Tech, 2026-07-10)
    High-performance computing (HPC) applications are increasingly dominated by data-intensive workloads that place new demands on storage subsystems. These workloads frequently diverge from traditional assumptions about I/O behavior, exhibiting irregular, small, and read-heavy access patterns that challenge the performance of distributed file systems. This dissertation explores empirical and predictive approaches to optimizing I/O performance in distributed object storage systems, focusing on workload-aware configuration strategies. This work systematically examines bandwidth and latency variations in the Ceph File System, built on top of the Ceph Distributed Object Store, due to diverse I/O patterns towards formalizing a data-driven relationship between latency, bandwidth, I/O size, and block-to-file size ratio. Our findings reveal that variations in system and application-level parameters can lead to substantial performance shifts, motivating the need for automated configuration support. We train a model for predicting bandwidth and estimating latency in similar workloads, providing storage administrators and application developers with actionable tuning recommendations, offering valuable insights into optimizing HPC storage I/O. Next, this dissertation focuses on the file striping strategy in CephFS, which utilizes a Raid-0 like policy to stripe incoming data over object(s) which are then mapped to the OSD(s). Currently, this striping strategy is determined at the time of cluster deployment and remains unchanged. We assess the impact of different stripe configurations by varying the object size, stripe count and stripe unit, on I/O performance across diverse workload types and access patterns. Our results highlight performance regimes where careful tuning yields significant gains, and expose inefficiencies in static, one-size-fits-all striping strategies. Overall, the focus of this dissertation is on the practical tuning of object storage systems grounded in empirical and predictive techniques enabling more responsive and efficient I/O performance in data-centric HPC workloads.
  • Deep Learning-Based Computational Tools for Microbial Protein Sequence Annotation
    Emon, Muhit Islam (Virginia Tech, 2026-07-09)
    Microbial protein annotation remains a fundamental challenge in understanding pathogen biology, host–microbe interactions, and the molecular basis of clinically important traits such as antimicrobial resistance and virulence. Traditional homology-based and manual feature engineering–driven approaches often struggle to generalize across diverse and rapidly evolving protein families, leaving a substantial fraction of microbial proteins poorly characterized. Recent advances in deep learning provide new opportunities to capture complex sequence and structural patterns beyond conventional techniques. This dissertation develops deep learning–based frameworks for microbial protein annotation across three biologically important and complementary contexts spanning bacteria, phages, and fungi. In Chapter 2, we introduce DeepMRP, a model for predicting bacterial metal resistance proteins, genetic elements that allow bacteria to withstand metal-based antimicrobials while also promoting the co-selection and persistence of antibiotic resistance across microbial communities. DeepMRP overcomes the limitations of fixed-threshold best-hit homology methods by leveraging bit score–based similarity distributions, enabling more robust identification of distantly related sequences. Chapter 3 focuses on phage structural proteins, which are essential for host recognition and phage–host interactions. We propose DeePSP-GIN, a structure-aware graph neural network that integrates predicted protein 3D structures with protein language model embeddings to capture both spatial and sequential dependencies, improving the identification and classification of phage structural proteins. In Chapter 4, we address fungal virulence factor proteins, key determinants of fungal pathogenicity and host invasion, by constructing a curated database and developing DeepFVFP, a protein language model–based framework with a convolutional neural network that captures biologically meaningful sequence patterns and enables interpretation of model predictions to uncover signals associated with virulence. Collectively, this work demonstrates that integrating sequence, structural, and learned representations through deep learning provides a powerful and flexible framework for microbial protein annotation. The approaches developed in this dissertation improve predictive performance and offer the potential to be extended to other functionally important protein classes, ultimately contributing to a deeper understanding of microbial systems and their roles in health and disease.
  • Multi-Stage Modeling with Gaussian Processes
    Flowers, Anna Rabren (Virginia Tech, 2026-07-09)
    Gaussian processes (GPs) furnish accurate nonlinear predictions with well-calibrated uncertainty. However, GPs alone are not always flexible enough to model complex real-world phenomena. To remedy this issue, it is beneficial to chain together multiple models. These so-called multi-stage models fit models in sequence, using output from one model as training data in the next model. I utilize multi-stage modeling with GPs to achieve two tasks. First, I improve GP prediction accuracy for data from processes with sudden changes, or "jumps,"' in the output variable by creating a new cluster-based (latent) feature and adding it to the input matrix. Then, I fit a GP to understand the circumstances under which a machine learning (ML) model performs best, and use that GP to propose the best circumstances in which to make future data acquisitions. I do this by treating the composition of metadata and the performance of the ML model, respectively, as the inputs and output of a GP. I vet both methods on a selection of real and synthetic benchmark examples from the recent literature.
  • Multilayer Liquid Metal Composite Circuits and their Fabrication for Soft Electronics
    Wilcox, Brittan Thomas (Virginia Tech, 2026-07-02)
    Next-generation wearable electronics demand flexible and stretchable circuity while retaining electrical functionality. Soft, liquid metal (LM) circuits show promise to achieve these properties. However, challenges exist in fabricating and integrating LM circuits into complex, multilayer designs such as adhesion to rigid components, the assembly of multiple circuit layers, and integration with surfaces of interest such as textiles. In this research, we will address the question: How can we control the assembly of LM within polymers and textiles and how do these architectures impact electrical, electromechanical, and mechanical properties within multilayer and textile-integrated soft circuits? This work will provide new methods and characterization of process parameters and material aspects to control the fabrication of multilayer LM circuits and the effects of adding bonded textiles for clothing-integrated wearable electronics. It will include investigation into how a textile layer impacts the electromechanics of LM circuits, how LM polymer systems perform underwater, how the movement of LM microdroplets in resin can aid in the assembly of multilayer soft circuits, and how LM composite structures are developed in lithography. By understanding these processes and material parameters, we will provide new understanding that can help realize wearable electronics with a combination of comfort and functionality that exceeds previous devices.
  • Companion dog nutrition epidemiology
    O'Brien, Janice Susan (Virginia Tech, 2026-07-08)
    Dr. Jan Sargeant once said, "the difference between statisticians and epidemiologists is that statisticians are concerned with random error, while epidemiologists are concerned with systematic error." The purpose of this thesis is to identify common causes of systematic error in canine nutrition research, especially studies conducted in large populations of dogs, to exemplify such examples through several analyses, and to provide contrasting examples of how to conduct population-based studies which minimize bias, and to further describe best practices specific to dietary exposures and health outcomes research in companion dogs. Owner factors are identified as likely confounders for diet and health outcomes, which are difficult to control for outside of randomized dietary assignment. Reverse-causation bias is also demonstrated to be a significant factor limiting the validity of cross-sectional analyses conducted here, therefore repeated measures analyses should be strongly considered in the field of companion dog nutrition epidemiology.
  • Differential expression of the mitochondrial calcium uniporter in hippocampal neurons attunes mitochondrial calcium handling and bioenergetics
    Cawley, Mikel Leann (Virginia Tech, 2026-07-08)
    Mitochondria are precisely positioned throughout neuronal axons and dendrites to meet unique energy demands across functionally distinct compartments. Dendritic mitochondria can be molecularly, structurally, and functionally distinct depending on the neuron type, supporting the diversity of synaptic functions and connectivity patterns across brain areas. A better understanding of how heterogeneous properties, such as mitochondrial morphology, dynamics, and function, converge to support cell- and compartment-specific metabolic demands, is crucial for elucidating how mitochondrial heterogeneity affects intact circuits. We found that the mitochondrial calcium uniporter (MCU) is endogenously enriched in the distal apical dendrites of CA2, precisely where CA2 neurons receive entorhinal cortical inputs carrying social information. Mitochondrial calcium uptake is tightly coupled with neuronal activity and stimulates the TCA cycle by activating calcium-dependent enzymes that are important for oxidative phosphorylation. Given the substantial energetic demand on neuronal mitochondria, we hypothesized that the enrichment of MCU in distal dendrites indicated mitochondria with a greater capacity to meet energy demand. To determine whether mitochondria enriched with MCU exhibited enhanced bioenergetic properties, we assessed whether MCU overexpression increased mitochondrial calcium uptake and respiration. We report enhanced mitochondrial calcium uptake rates in hippocampal mitochondria overexpressing MCU without increased sensitivity to calcium overload. Moreover, we reveal that the efficiency of mitochondrial respiration scales proportionally with bioenergetic demand. These findings suggest that differential expression of MCU can modulate mitochondrial calcium uptake and respiration under increasing demand and that mitochondria in CA2 distal dendrites have faster rates of calcium uptake without increased sensitivity to overload and synthesize ATP most efficiently under high energetic demand. These findings help expand our understanding of how mitochondrial adaptations support differences in neuronal bioenergetics, which may be critical for understanding cell-type-specific vulnerability to mitochondrial dysfunction in the brain.
  • Assessing crop yield and management yield zones using satellite imagery at the field scale
    Rathore, Jitender (Virginia Tech, 2026-07-08)
    Soybeans and corn are key crops in the United States, with many farms producing both. Predicting yield and delineating yield zones are crucial for nutrient management, guiding decisions to optimize production and increase yield. Accurate yield predictions are challenging due to variability in topography, soils, and management. We evaluated these factors using three (Edmunds, Hamlin, and Miner) South Dakota field trials and a set of six satellite images taken at different crop growth stages. Using Leave-One-Field-Out cross-validation, we trained an XGBoost machine learning model to predict yields and create spatial maps. Vegetation indices (VIs) from the R4/R5 stage showed the strongest correlation with soybean yield, especially Normalized Difference Vegetation Index (NDVI), Renormalized Difference Vegetation Index (RDVI), and Difference Vegetation Index (DVI), reaching 0.5-0.7 correlations. Soil Indices (SIs) had moderate correlation, with Coloration Index (CI) and Redness Index (RI) near 0.59-60 at Edmunds; Hamlin showed minimal soil-yield relationships. Topography showed a weak correlation, with elevation highest in 2019 (Pearson's R = 0.10) and 2021 (Pearson's R = 0.38). The best model at Edmunds (2021) had R2=0.54, RMSE=574 kg ha-1, and MAE=456 kg ha-1. Hamlin's performance was moderate; Miner performed poorly. Shapley Additive explanations (SHAP) analysis identified key predictors such as the DVI, RDVI, Green Chlorophyll Index (GCI), Triangular Greenness Index (TGI), and Green Normalized Difference Vegetation Index (GNDVI). In contrast, NDVI, Soil Adjusted Vegetation Index (SAVI), and Modified Soil Adjusted Vegetation Index 2 (MSAVI2) were deemed less significant. Combining topography, SIs, and VIs did not improve results. Focused on predicting high- and low-yielding zones using random forests (RF) and VIs, given that methods exist for delineating yield zones in crops. We hypothesized that integrating multi-year yield data, red-edge VIs, SIs, and topographic features (elevation, and slope) into a random forest (RF model) would improve soybean yield zone prediction. Predictors were tested across a range of Growing Degree Days (GDDs). Using Sentinel-2 satellite imagery from 2019 and 2021, twelve images were selected in a time series. An 80:20 split with k-5 cross-validation was used to train and test the RF model, evaluated by metrics like precision, F1-score, AUC-ROC, and accuracy. Error metrics were > 0.93, with best performance at GDD 1,080-1,300 and 1,300-1,500. Spatial maps were created to indicate yield potential expected under various weather conditions. To improve on the limitations of two-season data and to account for spatial-temporal variation, we collected historical yield data from Port Royal, VA, and delineated yield zones using clustering and spatial autocorrelation. We developed a workflow to clean, validate, and analyze multi-year data across crop rotations. Yield data from 78 site-years over 13 fields from 2018-2023 were filtered, aggregated into 10 m × 10 m grids, and spatially analyzed. Temporal stability facilitated classification of zones as high-stable, medium-stable, low-stable, or unstable, indicating long-term yield consistency. Unsupervised clustering enabled us to delineate management zones based on yield trends. We then compared average yields within zones and profitability potential, revealing distinct patterns with implications for site-specific management. Our results show that validated spatiotemporal yield maps effectively support zone classification and long-term decision-making for sustainable crop production.