Doctoral Dissertations

Permanent URI for this collection

Browse

Recent Submissions

Now showing 1 - 20 of 18544
  • Approaches, analyses, and applications of bat acoustic analyses
    Isenhour, Zackary William (Virginia Tech, 2026-08-14)
    Bat populations across eastern North America have experienced substantial declines due to white-nose syndrome, habitat alteration, and ongoing forest management pressures, increasing the need for accurate and locally informed approaches to monitoring and conservation. Acoustic surveys have become one of the most widely used tools for assessing bat distribution and habitat use; however, differences in sampling effort, automated call identification filtering, and statistical modeling approaches may substantially influence ecological inference and management recommendations. In this dissertation, I evaluated landscape-scale habitat associations, temporal activity patterns, and methodological considerations for acoustic monitoring of bats in managed forest landscapes of Pennsylvania and Maryland to provide evidence-based recommendations for improving acoustic survey design, data processing, and management decisions. First, I compared multiple approaches for modeling acoustic detections using generalized linear mixed models and eight acoustic data handling techniques for a rare species, the northern long-eared bat (Myotis septentrionalis), and a common species, the eastern red bat (Lasiurus borealis). I found that filtering acoustic detections using maximum likelihood estimate (MLE) thresholds from automated identification software improved prediction accuracy for rare species models relative to unfiltered data or highly conservative match-ratio thresholds. In contrast, the most conservative filtering approach produced the greatest predictive accuracy for eastern red bats, demonstrating that optimal acoustic processing approaches may differ among species according to detectability and prevalence. These findings suggest that managers should adopt species-specific acoustic data processing strategies rather than relying on a single filtering approach. For rare or difficult-to-detect species, such as federally listed bats, nightly maximum likelihood estimator (MLE)-retained detections provided the most reliable balance between minimizing false-positive identifications and retaining biologically meaningful observations. For common, readily detected species, more conservative filtering approaches may further improve predictive performance. Because species abundance and call characteristics vary geographically, practitioners should incorporate cross-validation during model development to identify the most appropriate data handling approach for their monitoring objectives. Second, I investigated multi-year landscape-scale habitat associations for nine bat species across Pennsylvania and Maryland. Several species exhibited distinct habitat relationships, including greater big brown bat (Eptesicus fuscus) activity along field edges and greater Lasiurus borealis activity in xeric conditions relative to hydric or mesic sites. Most species demonstrated relatively stable activity over the survey period despite low detections for several cave-obligate species affected by white-nose syndrome. Seasonal variation in acoustic activity suggested relationships with reproductive phenology (e.g., pregnancy, lactation, and juvenile volancy periods). Contrary to expectations, McNab's Landform Index was not strongly associated with activity for most species, including M. septentrionalis. Results also indicate that acoustic surveys conducted throughout the currently recommended survey window (May 15-August 15) better capture seasonal changes associated with pregnancy, lactation, and juvenile volancy than shorter sampling periods. Surveying both early and late portions of the season reduces the likelihood of incorrectly concluding that species are absent because life-stage-specific activity patterns were missed. Something USFWS should consider is requiring sampling in both early and late seasons to account for intraseasonal variation of activity. Finally, I assessed the influence of sampling effort and survey design on habitat association models by standardizing acoustic detector placement within a landscape-scale sampling grid. Results demonstrated that poorly matched survey effort and inconsistent detector deployment can generate uninformative or misleading habitat models, potentially biasing management recommendations and conservation actions. Collectively, this work demonstrates the importance of standardized sampling design, species-specific acoustic data filtering, and locally focused multi-year monitoring for improving the reliability of bat habitat association models and informing conservation and forest management decisions in eastern North America. I recommend that practitioners adopt MLE-retained detections for rare species, though further evaluations should be conducted to determine if these finding are appropriate range-wide. I also recommend to maintain acoustic studies across the full May 15-August 15 survey period (or beyond to capture true phenology) and increase sampling intensity where feasible to improve inference regarding species-environment relationships.
  • Elucidation of the Biosynthetic Pathway of the Iridoid Glycoside Monotropein and Its Association with Specialized Metabolism and Flavor in Blueberry
    Kaur, Ishveen (Virginia Tech, 2026-08-14)
    Blueberries (Vaccinium spp.) are valued for their health-promoting phytochemicals, including polyphenols and flavonoids, and some cultivars additionally accumulate monotropein, a bioactive iridoid glycoside with anti-inflammatory and neuroprotective properties. However, the genetic basis of monotropein accumulation and its potential relationship with fruit flavor metabolism are poorly understood. This study combined analytical chemistry, comparative genomics, transcriptomics, and metabolomics to investigate monotropein biosynthesis and its metabolic consequences in blueberry. Chapter 1: Optimization of protocol for accurate detection and quantification of monotropein (iridoid glycoside) from blueberry tissues Chapter 2: Identification of the iridoid glycoside monotropein biosynthetic pathway in blueberry using computational and molecular biology Chapter 3: Characterization of the variation in monotropein levels and its association with specialized metabolites in blueberry genotypes Blueberries are an economically important fruit crop rich in bioactive compounds like polyphenols and flavonoids. Blueberries have recently been reported to contain monotropein. However, methods to quantify monotropein levels in blueberries have not been optimized. To address this gap, we have developed and optimized an analytical method for monotropein extraction and quantification using liquid chromatography tandem mass spectrometry (LC-MS/MS). Different extraction strategies were compared, including variations in temperature, time, and ultrasonication treatments. Optimal extraction was achieved by heating samples to 60° C for 15 minutes in methanol. The method had high percent recovery (89-110% intraday; 91-108% interday) and good repeatability (1.17-2.15% relative standard deviation (RSD) intraday; 4.68-7.16% RSD interday). This protocol was then applied to 28 blueberry cultivars, 14 of which had not been previously analyzed for monotropein levels. Monotropein levels ranged from 0-1807 ng/mg (DW), The developed method can be applied to future evaluations of monotropein in diverse blueberry cultivars. To understand why cultivated blueberry produces lower levels of monotropein we performed comparative transcriptomic analysis and biosynthetic gene discovery in a pair of cultivars that produced high and low levels of monotropein. In our study, liquid chromatography-mass spectrometry (LC-MS) analysis identified 'Titan' as a monotropein-positive (M⁺) cultivar and 'Georgia Dawn' as a monotropein-negative (M⁻) cultivar. We then performed comparative orthology analysis to identify monotropein biosynthetic genes in blueberry using four different Vaccinium genotypes (Titan, Georgia Dawn, Draper (M-, V. corymbosum) and V. darrowii (M+, syn. V. ashei)) and the iridoid-producing plant Catharanthus roseus. Using orthology analysis and basic local alignment search tool (BLAST) we identified six of the core pathway genes in all Vaccinium genotypes. To further investigate genotype- and organ-specific regulation, transcriptomic and co-expression analysis was performed and correlated with metabolite data. Gene expression patterns showed no correlation with monotropein accumulation, suggesting that transcriptional regulation alone does not explain the observed metabolic differences. Moreover, gene cluster analysis further conveyed that the none of the genes in a monotropein biosynthetic pathway clustered together. We also identified three candidate oxidase genes that perform the final tailoring step in monotropein biosynthesis. Ongoing work focuses on functional validation of GES from M⁺ and M⁻ genotypes using heterologous expression systems. Enzymatic assays are needed to confirm the role of GES in determining monotropein biosynthesis. Finally, to determine whether increasing monotropein content could alter fruit quality, untargeted metabolomic profiling was conducted using LC-MS and GC-MS to identify different which are changing along with monotropein in different genotypes under study. We hypothesize that enhanced iridoid biosynthesis may also affect other metabolites as well in these genotypes. Monotropein abundance was positively associated with numerous flavonoids and terpenoid flavor compounds. These findings indicate different specialized metabolites including flavor metabolites might change in genotypes with varying amounts of monotropein. These changes can also affect consumer acceptance and marketability.
  • Behavioral Responses to Risk: Effects on Household Financial Decisions, International Portfolio Management, and Output Volatility
    Bandyopadhyay, Priyambada (Virginia Tech, 2026-08-13)
    This dissertation explores how behavioral responses to risk and uncertainty influence the financial decisions of individual households and financial entities involved in allocating external (foreign) capital to domestic investments. The analysis proceeds in two parts: the first chapter investigates how various trait types affect individual financial choices, while the second and third chapters examine how risk perception impacts resource allocation at the aggregate level. These allocation decisions have significant implications for output volatility. The first chapter draws on a large body of experimental psychology evidence indicating that smokers tend to be more risk-tolerant, impatient, and impulsive than non-smokers. As a result, smokers are inclined to make different economic decisions than non-smokers, which often manifests in financial and labor market behaviors. What is missing from this literature is a third category—quitters. The behaviors of this group may be particularly interesting because quitting smoking is a significant challenge, and overcoming this habit may involve unique behavioral and psychological traits that can manifest in actions and outcomes. In this chapter, I use data from the NLSY79 to classify individuals into three groups—smokers, non-smokers, and quitters—and examine how these groups differ in their financial decisions, with potential implications for current and future access to credit. I find that, compared to smokers, the rates of missed payments and bankruptcy are lower among quitters and non-smokers. Notably, in payment habits, quitters appear to be more prudent than non-smokers. The second chapter of the dissertation presents a theoretical model in which financial intermediaries play a central role. These intermediaries receive funds from domestic and foreign lenders and act to maximize depositors' returns. Because a sudden withdrawal of foreign funds (capital flight) reduces returns, intermediaries tend to offset the risk of capital flight by choosing a riskier portfolio, thereby increasing output volatility and lowering growth. This implied relationship sheds light on the literature linking actual capital outflows to macroeconomic volatility, as seen during the Mexican tequila crisis, the Asian financial crisis, and the Russian defaults. The chapter's contribution is to suggest that an increased perceived risk of capital outflow can also lead to higher output volatility and lower growth, even without capital actually flowing out of the country. To test this claim, we construct a cross-country measure from the IMF's Annual Report on Exchange Arrangements and Exchange Restrictions (AREAER) that captures the ease with which foreign capital can be withdrawn from a country. The data suggest that, after accounting for actual capital movements and other factors that influence output volatility, countries that allow unrestricted capital outflows tend to experience higher output volatility. The risk of capital flight from a country depends, unarguably, on how easily capital can be withdrawn. However, in practice, the severity of the threat may also depend on the volume of capital that could be lost due to market disruptions. The third chapter of the dissertation offers more refined support for the theory by treating the United States as the primary source of disruption. In particular, the chapter uses information on cross-country bilateral equity holdings from the IMF's Coordinated Portfolio Investment Survey to construct an index that measures both direct exposure to U.S. capital and indirect exposure to foreign capital from other countries whose capital flows co-move with U.S. flows. The results show that even after controlling for the volatility of realized flows, countries with higher exposure to the U.S. portfolio experience greater output volatility. This pattern is particularly evident among emerging and developing economies.
  • Aircraft Autonomous Contingency Landing Planning
    Tekaslan, Huseyin Emre (Virginia Tech, 2026-08-13)
    This dissertation develops an Assured Contingency Landing Management (ACLM) framework for time-critical contingency management of wing-lift-capable aircraft, enabling safe flight by minimizing risk to surrounding people and property while maintaining well-clear separation from nearby aircraft. Safe recovery from in-flight emergencies remains one of the most demanding challenges in modern aviation because contingency decisions must be made under severe time pressure. Traditional safety procedures rely heavily on the judgment and coordination of pilots and air traffic control (ATC). However, as airspace becomes increasingly congested and operational complexity increases, human cognitive limits and communication latency become critical bottlenecks. In dense urban environments, even seconds of delay can determine whether a feasible landing site remains reachable. The growing presence of Advanced Air Mobility (AAM) further amplifies these challenges by introducing large volumes of low-altitude air traffic near populated regions. To address these challenges, the proposed ACLM framework integrates landing site identification, trajectory generation, conflict resolution, and datalink communication for coordination into a unified real-time architecture. Its overarching objective is to advance the safety, reliability, and resilience of future airspace operations by jointly considering vehicle dynamics, airspace constraints, environmental conditions, nearby traffic, and risk to overflown populations within a data-driven, systems-level computational framework. This research advances both the theoretical foundations and practical implementation of ACLM through five major developments. The first component introduces a Dubins-based emergency landing path planning with air–ground coordination through datalink, for which emergency scenarios where road landings become necessary when no safer alternative is available. An analytical formulation is developed for multi-domain coordination, defining when and how to safely halt vehicular traffic to reserve a landing segment on a roadway identified through both offline geographic datasets and real-time ground traffic data. Numerical experiments validate the feasibility, safety, and timing requirements of this coordinated air–ground emergency response. The next development introduces foundations of the Gradient-Guided Search (GGS) algorithm, a discrete search-based risk-aware trajectory generation method that ensures feasible paths by guiding node expansion along local constraint gradients and discrete ground risk metric. This algorithm guarantees convergence to dynamically attainable trajectories. The planning results for varying initial emergency positions demonstrate improved risk avoidance over minimum-risk Dubins solutions, while maintaining low overhead suitable for onboard execution. The subsequent stage extends GGS into an airspace-aware planner that incorporates historical Automatic Dependent Surveillance–Broadcast (ADS-B) data and computational geometry to respectively model airspace density and proximity-based risk. By minimizing cumulative exposure to congested airspace, the planner reduces potential disruptions to nominal airspace operations. Next, the airspeed boundedness of GGS-based trajectories for engine-out flight is studied. A viability theory based framework capturing the coupling among guidance commands, wind, and altitude-to-airspeed exchange is developed for airspeed envelope protection. Admissible guidance commands are derived to guarantee forward-invariant airspeed and define safe maneuver primitives for embedding into the GGS planner. This integration eliminates the need for post-processing or online control modification under steady wind. Forward invariance under stochastic wind is further addressed through a hybrid extension that combines viability-certified primitives with online Control Barrier Function (CBF) based filtering. Numerical results demonstrate strict airspeed boundedness under steady wind and light-to-moderate gusts, while severe gusts require localized, short-lived CBF intervention without degrading path tracking performance. Finally, a real-time Plan-and-Avoid (PAA) framework is developed to coordinate cooperative multi-agent airspace operations around a declared priority trajectory. The priority trajectory represents an aircraft flight plan that must be preserved because of constrained maneuverability, an emergency, a mission-critical task, or assigned operational priority. The framework predicts uncertainty-aware well-clear separation violations with surrounding traffic and, when the priority trajectory alone cannot maintain separation, generates vehicle-constrained unilateral advisories that modify nearby aircraft trajectories to maintain well-clear separation for all traffic. Although applicable to any declared priority trajectory, the framework is demonstrated for emergency landing. The Plan component generates priority trajectories using the GGS method, extended to account for known multi-agent intent. PAA then identifies nearby aircraft that pass too close to the priority trajectory and issues Avoid resolution advisories to these cooperative aircraft. The framework is evaluated using real-world ADS-B traffic from the Washington, D.C., airspace. PAA generates feasible cooperative advisories for all identified conflict encounters and achieves real-time performance on a personal computer, including priority trajectory planning, advisory generation, and a 1 s two-way datalink delay. Most generated advisories also satisfy the RTCA DO-365 Detect-and-Avoid temporal threshold. These results demonstrate low-latency coordination that preserves priority trajectories while maintaining well-clear separation through automated real-time advisory generation. Overall, this work improves the foundation of ACLM by combining vehicle dynamics, environmental context, and coordination into an autonomy stack. The presented algorithms are scalable and contribute toward resilient, self-managing airspace operations that advance existing procedures in terms of both safety and response time. Future research will focus on adaptive replanning under uncertainty and the extension of the presented ACLM principles to rotorcraft and different failure modes, ultimately contributing to a scalable infrastructure for autonomous aircraft systems and their adoption for safe airspace operations.
  • Modulating the Earth's Magnetic Field for Communication in RF Denied Environments
    Chinnakkagari, Shashank Reddy (Virginia Tech, 2026-08-13)
    Low-frequency magnetic-field communication is attractive for radio-frequency (RF)-denied environments, including underwater, underground, and through-the-earth (TTE) communication, because electromagnetic fields at ultra-low frequency (ULF) and very-low frequency (VLF) experience substantially lower attenuation in conductive media than conventional RF signals. However, the long wavelengths associated with ULF/VLF operation make conventional antennas inefficient, physically large, and power intensive. Recent approaches, including mechanically based antennas and rotating permanent-magnet transmitters, reduce transmitter size but introduce moving parts, mechanical reliability concerns, and scaling limitations. This dissertation investigates controlled permeability modulation as a non-mechanical method for generating low-frequency magnetic signals. The central idea is that a static magnetic field can be converted into a time-varying magnetic-field signal by electrically varying the permeability of a nearby ferromagnetic structure. Two implementations of this principle are studied. The first uses a stationary permanent magnet as the static magnetic-field source. The magnetic flux from the permanent magnet is modulated by changing the effective permeability of a surrounding Metglas shield using an electrically driven control coil. The transmitter is analyzed under both high-power and low-power excitation, and the dependence of the modulation behavior on shield saturation, magnet orientation, control-coil excitation, and material properties is investigated through simulation and measurement. Analytical models are developed to describe flux redistribution in single-layer and multilayer ferromagnetic shielding structures containing an internal magnetic dipole source. These models relate magnetic shielding behavior to permeability, shield geometry, interlayer spacing, and magnet dimensions. A multilayer Metglas shield is then introduced to improve modulation depth while reducing shield weight and control power. Prototype transmitters are fabricated and experimentally characterized using three-axis air-core receiver coils. The second and the main contribution eliminates the permanent magnet and uses the ambient geomagnetic field as the static magnetic bias source. A ferromagnetic structure aligned with the local magnetic north–south direction perturbs the local geomagnetic field, and modulation is achieved by varying the permeability of the structure with a control coil. Measurements verify the orientation dependence of the received signal and show second-harmonic modulation under sinusoidal excitation, consistent with cyclic permeability modulation of the ferromagnetic material. These results demonstrate the feasibility of compact, non-mechanical ULF magnetic-field transmitters based on controlled modulation of a static magnetic field, including the ambient geomagnetic field. The magnetic-field generation concepts developed in this dissertation also suggest broader applications beyond communication, including rotating-field transcranial magnetic stimulation (TMS) systems, which are treated as a related application in the appendix and are the subject of separate publications.
  • Quantum Online Learning
    Zheng, Alice (Virginia Tech, 2026-08-12)
    Online learning is a framework for interfacing with environments whose structure is revealed through interactions; quantum information theory is a means of representing data and its processing consistent with modern understanding of physics. Their union is symbiotic---sequential processing is light on precious quantum resources, and quantum algorithms can surpass known bounds of classical algorithms---their combination being an area of independent interest. This thesis aims to exemplify the above point via five case studies on tomography, change detection, and function evaluation. In the process of deriving efficient tomography procedures, prior work has shown that quantum states can be learned in online adversarial environments. We extend this notion to subsets of positive semidefinite operators, including general quantum objects such as states, measurements, channels, strategies, co-strategies, and Gram matrices. We tailor a regularized follow-the-leader algorithm and achieve sublinear regret. As a byproduct, we show a generalization of Pinsker's inequality that accounts for differences in trace, its proof method more broadly applicable to general divergences. We further extend online learnability to objects in arbitrary physical theories. We derive mappings between Euclidean Jordan algebras and generalized probabilistic theories, using methods for the former to show that states in the latter are learnable. Specifically, we develop a projective version of the symmetric cone multiplicative weights update algorithm that achieves sublinear regret when learning states of generalized probabilistic theories. We consider the task of detecting changes in sequences of unknown quantum states with a constraint of no false positives, developing efficient and optimal online quantum algorithms. These include online algorithms with small amounts of quantum memory based on the swap test, and optimal algorithms projecting onto the symmetric subspace. We find circuit constructions for these algorithms, analyze their complexity, and show their equivalence under certain memory limitations. The proposed algorithms are applicable to quality control and anomaly detection in quantum devices, and may be of independent interest for identity testing and optimal cloning. We elaborate on the quality control aspect of changepoint methods by developing procedures for aiding the calibration of near-term quantum devices. These consist of algorithms for certification and changepoint detection in Hamiltonian dynamics, which enable continuous monitoring and trigger recalibration procedures. Using a cumulative sum procedure, we attain an asymptotically-optimal scaling of the average times to detection versus false positive, dependent primarily on Hamiltonian norm bounds. Lastly, we quantify the effect of memory limits on Boolean function evaluation of quantum sequences---a task applicable to distributed sensing and quantum reading of classical memories, among others. We show that minimum-error function evaluation in the memoryless regime is equivalent to the so-called pretty good measurement and hence shares its performance guarantee. Additionally, we characterize the precise set of functions on which this strategy performs optimally---affine functions. The above studies reveal insights about learnability as a universal property and the tradeoffs between performance and resource requirements in online settings, informing the choice of quantum settings best-suited for online learning. We conclude by summarizing these takeaways, along with listing remaining open questions.
  • Exploration of Dynamic Structure-Process-Property Relationships in Vitrimer-Like Materials and Multimodal Polymer-Clay Composites
    Luster, Larry E. (Virginia Tech, 2026-08-12)
    When designing and modeling composite materials, it is common to make simplifying assumptions regarding intermolecular interactions so that the composite can be considered as a homogeneous material with predictable behavior. This approach is acceptable for traditional composite processing and in any case where the components of the composite have uniform, static structures after they are incorporated into the bulk material; however, real materials for engineering applications rarely fit these idealized models and the components have specific, non-negligible interactions and transient topologies. This body of work examines two such materials: vitrimeric thermoplastic polymers and polymers reinforced with layered aluminosilicates; specifically, we 1) attempt to deconvolute contributions of bond exchange kinetics and segmental relaxation in a covalent adaptive network comprised of a poly(methyl-methacrylate)-poly(hydroxy-ethyl-methacrylate) copolymer (PMMA-PHEMA) crosslinked with dynamic aromatic disulfide bonds and 2) examine the role of in-situ dehydration on intercalation and exfoliation of montmorillonite agglomerates into nanoplatelets during melt-extrusion of a polyethylene terephthalate glycol (PETG)-montmorillonite-zeolite composite. In both studies, we challenge the fundamental assumptions used to simplify kinetic, thermodynamic, and transport properties of the material and utilize bulk rheological measurements to extrapolate mechanistic understandings of the material behavior that can be exploited in process design to yield desirable properties and morphologies in the end-use material.
  • Exploring the Potential Safety Impact of Automated Driving Systems Using Naturalistic Data
    Herbers, Eileen Mary (Virginia Tech, 2026-08-12)
    Automated driving systems (ADS) have the possibility to remove humans from the primary driving task, which has the potential to eliminate all crashes due to human driver error. However, a large component preventing the wide-scale adoption of ADS is the ability to confidently determine when they are safe enough to deploy and achieve the associated societal acceptance. The question of "how safe is safe enough?" has consumed the industry for quite some time. Determining the threshold of "safe enough," however, has proven to be quite a complex problem. A common perspective suggests that as long as ADS are safer than human drivers, then they are safe enough to deploy. However, this perspective introduces several challenges, including determining an appropriate human driver baseline (some drivers are safer than others), defining the conditions and scale of comparison, and considering whether society is willing to accept current roadway risks, which result in over forty thousand fatalities each year [1]. In these assessments, it is relevant to note that most driving is routine. Thus, modest ADS deployments are rarely exposed to the complex scenarios that the 233 million licensed human drivers encounter across the United States [2]. In these rare situations, human drivers might outperform ADS unless we identify methods to characterize these events and better understand what we are expecting robotic drivers to achieve. Naturalistic driving data, which represents large-scale in situ data collections of everyday driving, are used within this dissertation to identify and validate events and scenarios that might challenge the capabilities of ADS, indicating areas where further development and validation may be required to achieve acceptable levels of safety. Specifically, this dissertation analyzes nearly three thousand safety-critical events (SCEs) that involve crashes and near-crashes in a variety of scenarios. Specific events are selected to evaluate conditions under which ADS may have difficulty navigating the situation correctly. The analysis focuses on surprise events, including scenarios in which line of sight (LOS) perception systems are obstructed and constrain system response, as well as events involving unpredictable behavior by other road users in which the subject driver is not at fault. This analysis suggests that ADS may not perform as expected in blind turns and hills, mixed-speed traffic, lane-change events with other vehicles around, scenarios with significant occlusion at high speeds, and in scenarios in which pedestrians are visually occluded. By investigating SCE scenarios that are believed to be challenging for ADS to navigate, this research shows that using a small set of naturalistic data has the potential to convey important information to wide-scale ADS deployment that simulation or closed-track testing based on contrived scenarios simply cannot achieve. Near-crash and crash-relevant events are especially crucial for fully understanding the complex driving task and should be studied further to completely assess ADS safety. Additionally, human drivers are generally good at performing evasive maneuvers that require a complex understanding of the surrounding environment. Such near-crash situations necessitate an intricate sequence of perception and response for ADS, which may not be fully understood based on data from the limited ADS deployments performed to date. This dissertation aims to inform selection of ADS testing scenarios, catalogue edge cases which challenge ADS and limit their ability to provide safe transport, and to identify the opportunity to improve ADS performance in such situations through inclusion Vehicle-to-Everything (V2X) communications which enable perception beyond LOS sensing.
  • Robust Functional-Input Hypothesis Testing and Multi-Task Regression on Mixed-Type Graphs for High-Dimensional, Complex-Structured Data
    Chen, Mengkun (Virginia Tech, 2026-08-10)
    High-dimensional, complex-structured data have become increasingly prevalent across modern scientific disciplines such as social sciences, omics, and medical imaging. As such data continue to proliferate, developing advanced statistical methodologies for scalable estimation and inference is essential for extracting meaningful insights and driving scientific discovery. This dissertation introduces three robust, flexible, and computationally efficient nonparametric methods for functional data analysis: (1) a flexible test for detecting unknown functional departures under generalized functional regression for biomedical group discrimination; (2) a robust functional-input kernel-machine-based test for identifying brain networks associated with health outcomes; and (3) a joint functional regression on mixed-type functional graphs for dynamic network modeling in brain imaging data. For the first method, we introduce a hybrid Frequentist-Bayesian hypothesis testing procedure that combines a Bayes factor and a score-type statistic to detect nonlinear and nonparametric effects of functional predictors on scalar responses. This approach avoids restrictive parametric assumptions and explicit likelihood estimation, offering a practical and flexible tool for biomedical group discrimination. For the second method, we develop a kernel-machine-based quantile regression test to identify associations between high-dimensional, correlated functional predictors and heavy-tailed or skewed responses. This robust approach accommodates complex interactions and successfully identifies autism spectrum disorder (ASD)-related brain networks supported by neuroscience evidence. For the third method, we propose a unified joint modeling framework that simultaneously selects functional predictors embedded in a time-varying functional graph and estimates dynamic network structures without prior information. The model combines Bayesian hierarchical modeling by developing a computationally efficient marginal expected integrated penalized likelihood maximization (ME-IPLM) algorithm, enhancing variable and graph selection accuracy, interpretability, and prediction performance. Together, these three projects provide scalable, interpretable, and theoretically supported frameworks for analyzing high-dimensional functional data, advancing group discrimination, association detection, and dynamic network interpretation across diverse scientific domains.
  • Mechanisms of Nutrient Decline in Soybean under Elevated CO₂: A Physiological and Transcriptomic Investigation
    Kaur, Ravneet (Virginia Tech, 2026-08-10)
    Rising atmospheric carbon dioxide (CO₂) can stimulate photosynthesis, biomass production, and yield in C₃ crops, but these responses are often accompanied by lower seed mineral concentrations. The physiological and molecular processes underlying this decline remain poorly resolved in soybean (Glycine max L. Merr.), an important source of protein and minerals. This dissertation integrated physiological measurements, ionomic analysis, transcriptomics, and canopy-level seed sampling to determine how carbon assimilation, water use, nutrient uptake and allocation, gene expression, and canopy position contribute to soybean seed nutrient responses under elevated CO₂ (eCO₂). Soybean cultivars with contrasting yield, nutrient, and water-use responses were grown under ambient and elevated CO₂ in open-top or controlled-environment chambers. Across experiments, eCO₂ increased carbon assimilation, aboveground biomass, seed number, and seed yield in responsive cultivars, while root biomass and nutrient uptake did not increase proportionally with reproductive growth. Mature seed concentrations of several macro- and micronutrients declined, indicating that carbon-driven seed production exceeded the capacity of plants to acquire and allocate minerals to developing seeds. Transcriptomic analysis of leaves, roots, and seeds at the beginning seed stage showed that organ identity accounted for most variation in gene expression, with roots exhibiting the largest response to eCO₂. Leaf expression patterns were associated with photosynthesis, carbohydrate metabolism, and ion transport. Root responses were associated with cell wall modification, oxidative functions, hormone metabolism, and nutrient transport, while seed responses were associated with carbohydrate metabolism and cell wall modification. Most differentially expressed genes were restricted to individual cultivar-by-organ combinations, showing that similar seed mineral outcomes were not preceded by one shared transcriptional response. Instead, the combined physiological and molecular data indicated a disconnect between carbon-driven growth and proportional nutrient uptake and allocation. This interpretation was further evaluated using cultivars selected for contrasting transpiration and yield responses. Elevated CO₂ reduced transpiration by 8.4% and stomatal conductance by 37.5%, but seed nutrient accumulation followed cultivar yield response rather than water-use phenotype. High-yielding cultivars accumulated more total minerals in seeds, although this increase did not keep pace with seed production, whereas cultivars with little or negative yield responses accumulated fewer seed minerals. Canopy position also strongly influenced nutrient distribution, with middle-canopy seeds contributing the greatest seed mass and mineral accumulation regardless of CO₂ treatment. Together, these findings show that soybean seed nutrient decline under eCO₂ is driven primarily by carbon assimilation and reproductive growth outpacing nutrient acquisition and allocation, while reduced transpiration provides a less consistent explanation. The cultivar- and organ-specific responses identified here demonstrate that similar declines in seed mineral concentration can arise through different physiological and transcriptional relationships. This work provides a framework for identifying soybean traits and cultivars that maintain seed nutritional quality under future atmospheric conditions.
  • Defending the Throne: Leader Narcissism, Status Threat, and the Self-Regulatory Path to Abusive Supervision Intentions
    Conger, Joseph Zachary (Virginia Tech, 2026-08-10)
    This dissertation examined the relationships between unidimensional and tripartite narcissism, four mediating cognitive self-regulatory responses, and abusive supervision intentions. I expected unidimensional narcissism to have positive associations with the selected self-regulatory responses and abusive supervision intentions. I recruited 202 full-time workers with supervisory experience to partake in a between-person, paper-person vignette experiment. Participants imagined themselves in a story involving a follower they had to evaluate, then answered scales capturing cognitive self-regulation and their intent to employ abusive supervision. Participants were randomly assigned to a follower engaging in neutral or status-threatening behavior. Results from two-stage path analysis revealed that, in the neutral condition, unidimensional narcissism's impact on abusive supervision intentions was fully mediated by revenge motivations. In the threat condition, narcissism's impact on abusive supervision intentions was partially mediated by revenge motivations and moral disengagement. Threat condition did not significantly strengthen the relationships between narcissism and cognitive self-regulatory responses. A research question tested how tripartite narcissism related to self-regulation variables in the threat condition. Findings suggested self-important narcissism increased revenge motivations and moral disengagement, as well as decreased job performance ratings of the follower. Direct effect estimates revealed grandiose narcissism negatively predicted and self-important narcissism positively predicted abusive supervision intentions. Vulnerable narcissism had no significant impact on any endogenous variables. I conclude that, whether the threat is real or imagined, derogation of others is a key tool narcissists use to defend their status.
  • Building Trustworthy Machine Learning Systems for Security: From Federated Learning to Agentic AI
    Zhang, Chaoyu (Virginia Tech, 2026-08-07)
    Federated Learning (FL) has emerged as a promising paradigm for privacy-preserving collaborative machine learning, enabling distributed devices to jointly train models without sharing raw data. This dissertation addresses critical challenges in building trustworthy and secure machine learning systems across two complementary frontiers: FL systems for network and medical security, and anomaly detection for agentic AI. Together, the five technical chapters form a progression from improving FL's utility and resilience, through applying FL to network security, to defending FL from model inversion attacks in medical settings, and finally to detecting workflow-level anomalies in modern agentic AI systems. Chapter 2 tackles data quality heterogeneity in FL to improve global model utility. Noisy-labeled and imbalanced local data among clients can severely hinder training efficiency and model convergence. We propose a quality and fairness-aware client selection mechanism based on a novel Quality-of-Model (QoM) metric that evaluates client contribution without requiring access to model updates or gradients. Our approach prioritizes clients with high-quality contributions while ensuring diversity through a sortition-inspired randomized selection process, improving convergence speed and reducing communication overhead. Chapter 3 addresses Byzantine resilience to ensure FL trustworthiness from a system security perspective. FL's distributed nature forces the central server to blindly trust local training processes, making it vulnerable to model poisoning and data poisoning attacks from malicious participants. We propose a remote attestation-based approach that regains transparency into client-side training by verifying computation integrity via Trusted Execution Environment (TEE)-generated cryptographic attestation reports, enabling the server to reject malicious updates while preserving performance under non-IID data distributions. Building on these foundations, Chapter 4 demonstrates FL applied to network intrusion detection. We propose a geometric feature learning approach that projects network traffic into a compact, well-structured representation space, combining contrastive feature learning with H-Score optimization to maximize intra-class compactness and inter-class separability. The resulting federated intrusion detection system supports both anomaly detection and precise attack type identification, including zero-day threat exploration through entropy-based uncertainty analysis. Chapter 5 addresses a distinct and critical privacy threat in FL: model inversion attacks (MIAs) that allow a malicious server to reconstruct private training data from shared model updates. We identify that the reconstruction success of all known MIAs is fundamentally bounded by the local batch size relative to the model's first-layer leakage capacity. Building on this observation, we propose Aegis, a defense that synthesizes an auxiliary dataset to push the effective batch size beyond this capacity, collapsing the server's closed-form reconstructions without modifying the FL protocol or perturbing patient data. We demonstrate the approach on medical imaging datasets. Chapter 6 extends our security focus beyond FL to the rapidly growing domain of agentic AI. Modern agentic systems execute complex tasks through long-horizon workflows involving multi-agent coordination and tool invocation, creating a new risk surface where a single injected or erroneous step propagates through downstream dependencies. We present Skynet, a workflow-level anomaly detection framework that models multi-agent execution as directed workflow graphs and learns benign behavior jointly over semantic and structural dimensions. Trained exclusively on benign workflows, Skynet detects both adversarial manipulations and intrinsic execution failures under a single decision rule, naturally extending to zero-day anomalies. Together, these five chapters provide a comprehensive framework for building trustworthy machine learning systems: improving FL utility and trustworthiness, applying FL to network and medical security, and detecting anomalies in emerging agentic AI architectures, advancing the state of the art in machine learning security in adversarial and distributed environments.
  • Exploring Generative AI for Pedagogically Aligned Learning Experiences and Adapted Instructional Practices in Software Engineering Education
    Wang, Tianjia (Virginia Tech, 2026-08-05)
    In the era of generative artificial intelligence(AI), educators and students face both practical opportunities and pressing challenges for software engineering (SE) education. Generative AI powered by large language models is capable of completing complex software development tasks, becoming increasingly integrated into the development process, and is reshaping how students can learn to design, develop, and test software systems. With the advancement of generative AI, it is important to understand how students and educators perceive its role, benefits, and challenges in educational contexts. Also, many existing generative AI tools are adapted from general-purpose models without considering how they align with curriculum goals, cognitive development, or instructional strategies in SE education. In this dissertation, we present our examination of generative AI in SE education from: (1) understanding the perceptions, practices, and expectations of students and instructors regarding generative AI; (2) exploring how intelligent systems powered by generative AI can align with pedagogical goals and support active, collaborative and engaging learning environments; (3) investigating how generative AI could reshape SE knowledge areas to guide curriculum design and assessment. The findings show that generative AI can solve course problem sets with high accuracy, raising concerns about academic integrity, misinformation, and shallow learning; however, instructors remain open to its use when supported by clear policies and redesigned assessments. Students valued generative AI for providing real-time feedback, decomposing complex tasks, and reducing anxiety among those hesitant to ask questions, but identified limitations in emotional connection, nonverbal communication, and human-like interaction. In investigating how generative AI can effectively support student learning while addressing misuse and related issues, the studies on intelligent systems demonstrate that generative AI agents can simulate believable behaviors in learning environments, increasing students' perceived cognitive and social presence. The findings further indicate that generative AI is more effective in supporting student learning and engagement when learning objectives are embedded in system design and AI agents' behaviors are aligned with pedagogical goals. The AI-powered multi-agent system can provide structured, continuous, and interactive scaffolding to help students learn the software development life cycle, significantly improving students' learning gains, increasing task completion rates, and strengthening students' engagement. Finally, the dissertation provides empirical findings that indicate inconsistent topic coverage in undergraduate SE courses. Developers perceived GenAI as most effective for tasks in coding-related knowledge areas. At the same time, participants did not view GenAI as replacing foundational SE knowledge and identified a shift in the importance of the SE knowledge areas. Developers also emphasized the need to include AI-related knowledge and skills in SE education, suggesting that SE curricula should adapt toward preparing students to engineer, validate, and take responsibility for software development produced through human-AI collaboration.
  • Looking for New Particle Physics with Astrophysical Origin
    Gustafson, Robert Andrew (Virginia Tech, 2026-08-05)
    In this thesis, I explore the consequences of introducing new particles into astrophysical environments, and place constraints on these particles using available data. I first consider Heavy Neutral Leptons (a proposed particle which has important implications for neutrinos) in the context of atmospheric interactions, the Sun, and supernovae. I then turn my focus to various models of dark matter, considering the reach of both the terrestrial and astronomical observables. A special focus is given to cases where dark matter clusters around Supermassive Black Holes.
  • Automated Testing for Data-Intensive Scalable Computing
    Humayun, Ahmad (Virginia Tech, 2026-08-05)
    Data-Intensive Scalable Computing (DISC) systems such as Apache Spark, Hadoop, Flink, and Beam have become central to modern data processing. These systems enable developers to write applications as dataflows composed of user-defined functions operating over large and often unstructured datasets. However, testing such applications and the underlying frameworks remains a significant challenge. Conventional testing techniques struggle in this space due to the complexity of dataflow semantics, the scale and heterogeneity of input data, and the need to reason about operator interactions, schema variability, and program logic simultaneously. This thesis presents a set of techniques for extracting and leveraging rich, fine-grained properties that are unique to DISC workloads in order to inform automated input generation for effective testing across various levels of the data-centric software stack. The first thrust of this work focuses on generating inputs that can avoid trivial parsing errors and effectively exercise the deeper logic in DISC applications. By analyzing how code interacts with different parts of the input data, we test the code where it matters most instead of wasteful fuzzing cycles finding parsing issues. The second thrust explores how to create realistic and meaningful test data that reflects the structure and semantics of real-world inputs, while still achieving high coverage and fault detection. The final part of this work shifts focus to DISC frameworks, recognizing that applications are only as reliable as the systems they run on. By generating diverse dataflow programs, we systematically test internal components like optimizers, extending fuzzing to the entire DISC stack.
  • Explainable Machine Learning Framework for Accurate Reference-free Biological Data Deconvolution
    Du, Dongping (Virginia Tech, 2026-08-05)
    Bulk omics data captures mixed molecular signals from multicellular tissue compositions, posing significant challenges in accurate interpretation of biological changes over samples. Reference free deconvolution offers a flexible framework to estimate cell type proportions and specific expressions from bulk data, thus uncovering latent cellular and molecular architecture of tissue ecosystems without relying on predefined references. However, existing deconvolution methods are highly sensitive to several hidden confounders, including asymmetric gene expression, informative missingness, inter-cell-type imbalance, and deviation from identifiability conditions. These issues also propagate across preprocessing and modeling stages, collectively, leading to reduced accuracy of bulk deconvolution and downstream inference. In this dissertation, we present a comprehensive methodological framework with effective workflow for reference-free deconvolution of complex biological data. The proposed workflow integrates four key components spanning preprocessing, missing value imputation, structural correction, and discriminative deconvolution. First, we propose Cosbin, an iterative normalization strategy that identifies consistently expressed genes and removes asymmetrically differentially expressed genes, thereby preserving the geometric structure of the data. Second, we propose mechanism-integrated group-wise pre-imputation, which explicitly models multiple missingness mechanisms and preserves biologically informative missing patterns, particularly for marker-like genes. Third, we introduce iterative equilibration of cell-type expression profiles to correct inter cell-type asymmetry and improve the identifiability and accuracy of proportion estimation. Fourth, we apply and evaluate CAM3.0, an enhanced convex geometry-based unsupervised deconvolution algorithm, to estimate latent molecular archetypes and compositions on diverse omics data types from real biological bulk samples. To further improve deconvolution accuracy and efficiency particularly when constituent cell types are highly mixed or hardly separable, we also propose and develop an effective cosine similarity based discriminative analysis of mixtures method (csDAM), specifically to deconvolute highly mixed bulk expression data. Facilitated by the monotonic relationship between signature gene specificity and cosine similarity rank distribution, csDAM achieves highly accurate estimation of cell type proportions and specific expressions. Simulation and real-data based studies show both improved deconvolution performance and computational efficiency by csDAM compared to most relevant peer methods. In summary, the work in this dissertation addresses several major limitations of existing reference free deconvolution approaches by collectively integrating improved preprocessing, missing value imputation, structural correction, and discriminative modeling into a unified framework. Extensive simulations and real-data applications demonstrate improved performance in terms of stability, accuracy, and interpretability under challenging conditions, including high noise, complex missingness, and strong cellular heterogeneity. This study highlights the critical role of structural considerations in deconvolution and provides a scalable solution for extracting biologically meaningful latent features and archetypes from large-scale bulk omics data.
  • Computational Framework and Deep Learning for Cross-Species Comparison in Plant Transcriptomics and Imaging
    Chau, Tran Ngoc (Virginia Tech, 2026-08-05)
    Cross-species cell type analysis is fundamental to comparative plant biology, enabling the transfer of knowledge from well-studied model species such as Arabidopsis thaliana to agriculturally and medicinally important non-model species. Recent advances in single-cell RNA sequencing (scRNA-seq) have generated large-scale transcriptomic atlases across diverse plant species, yet extracting biological insight from these datasets, and comparing them across species, remains challenging due to data sparsity, dropout noise, limited marker gene information in non-model plants, limited ortholog detection across distant species, and the lack of integrated frameworks across data modalities. This dissertation develops computational frameworks for cross-species cell type mapping in plants, spanning both transcriptomic and imaging data, through a sequence of connected projects. First, we address the lack of reliable tools for identifying co-expressed gene modules and imputing missing values to improve data quality in plant scRNA-seq. We construct a benchmark that evaluates co-expression and imputation methods against a ground truth of promoter-reporter-validated native gene pairs, providing the community with a trusted foundation for downstream analyses such as identifying transcription factor target genes, nuclear pore complex gene associations, and functional gene modules. Second, we introduce the Orthologous Marker Gene (OMG) framework, which maps cell types across 15 plant species by identifying cell-type-specific marker genes that belong to conserved orthogroups. Rather than relying on one-to-one gene-level orthology, which is confounded by sequence divergence across distantly related species, OMG aggregates markers at the orthogroup level, enabling robust cross-species cell type correspondence. Then, we extend this framework with PLM-OMG, a protein language model-based approach for scalable orthogroup classification, enabling incremental inclusion of new species without recomputing existing orthogroups, substantially reducing computational overhead and improving cross-species cell type mapping. Finally, we extend cross-species cell type analysis to imaging by developing a hybrid pipeline for confocal root images across different species and developmental stages. The pipeline integrates Cellpose-SAM segmentation with a classifier that combines morpho-topological features and fine-tuned DINOv2 vision transformer embeddings, trained using gradient-boosted models and iterative refinement. Ambiguous cells are resolved using a vision-language model, demonstrating scalable automated cell type annotation across diverse plant species. Together, these contributions provide reusable computational frameworks for cross-species cell type mapping in plants, spanning transcriptomic and imaging modalities, and advancing comparative plant genomics at scale.
  • From Protest to Policy: An AI-Assisted Meta-Analytic Journey Through LGBTQ+ Change
    Cornett, Kelsi (Virginia Tech, 2026-07-31)
    This dissertation examines how activism, public opinion, and legal change interact in the advancement of LGBTQ+ rights. A human-in-the-loop, AI-assisted meta-analytic workflow employed ASReview LAB, Elicit AI, and GPT to support screening, data extraction, and validation. The final synthesis included 28 effects represented in 25 independent study matrices with a summed analytic sample size of N=1,613,593. Random-effects meta-analytic structural equation modeling (MASEM) was used to pool associations among the three constructs and compare competing conceptual models of sociopolitical change. Activism and public opinion were each positively associated with policy or legal change, whereas the direct association between activism and public opinion was small and nonsignificant. The strongest-fitting model was the policy-responsive mobilization model (Model 9; CFI= 1.00, TLI= 1.278, AIC= -1.98, BIC=-14.27). In this model, public opinion predicts policy or legal change (b = 0.186), which subsequently predicts activism (b = 0.187). However, the indirect effect was not statistically significant, so the findings do not confirm mediation or causality. The public opinion-policy relationship was the most stable across sensitivity and influence analyses. Nevertheless, the evidence base was highly heterogeneous, temporally constrained to the 21st century, concentrated in the Global North, and unevenly distributed across relationships. Overall, the findings suggest that legal reform may function as both an outcome of social movements and as a source of legitimacy and opportunity for renewed activism.
  • Explaining and Steering Embedding Projections for Visual Analytics
    Liu, Wei (Virginia Tech, 2026-07-29)
    Low-dimensional embedding projections are widely used in visual analytics, particularly for exploring large document collections. By arranging documents as points in a two-dimensional space, these projections help analysts identify clusters, outliers, separations, and relationships among documents. However, projection layouts are often difficult to interpret and control: users can observe where documents are positioned, but may not understand why spatial patterns appear or how to reshape the projection when the resulting layout does not align with their analytic goals. This dissertation frames embedding projections as interactive semantic workspaces and develops methods for explaining and steering them in visual analytics. First, it introduces gradient-based explanations that connect textual features to document positions in projection layouts, revealing how words influence spatial placement. Second, it presents context-aware natural-language explanations that combine document semantics with layout-derived spatial context to help users interpret documents, regions, and spatial patterns. Third, it moves from explanation to steering by introducing an LLM-augmented semantic steering approach, in which analysts express semantic intent through example groupings and reshape projections without retraining the underlying models. Finally, it develops a scalable prototype-based steering method that shifts LLM reasoning from individual items to group-level abstraction, making semantic steering practical for large embedding collections. Through quantitative evaluations, usage scenarios, case studies, and a user study, this dissertation demonstrates that embedding projections can be made more interpretable, controllable, and aligned with analytic goals. These contributions advance projection-based visual analytics toward interactive semantic workspaces that analysts can inspect, understand, and reshape.