Building Trustworthy Machine Learning Systems for Security: From Federated Learning to Agentic AI
| dc.contributor.author | Zhang, Chaoyu | en |
| dc.contributor.committeechair | Lou, Wenjing | en |
| dc.contributor.committeemember | Ramakrishnan, Narendran | en |
| dc.contributor.committeemember | Zhang, Ning | en |
| dc.contributor.committeemember | Ji, Bo | en |
| dc.contributor.committeemember | Hou, Yiwei Thomas | en |
| dc.contributor.department | Computer Science and#38; Applications | en |
| dc.date.accessioned | 2026-08-08T08:00:12Z | en |
| dc.date.available | 2026-08-08T08:00:12Z | en |
| dc.date.issued | 2026-08-07 | en |
| dc.description.abstract | Federated Learning (FL) has emerged as a promising paradigm for privacy-preserving collaborative machine learning, enabling distributed devices to jointly train models without sharing raw data. This dissertation addresses critical challenges in building trustworthy and secure machine learning systems across two complementary frontiers: FL systems for network and medical security, and anomaly detection for agentic AI. Together, the five technical chapters form a progression from improving FL's utility and resilience, through applying FL to network security, to defending FL from model inversion attacks in medical settings, and finally to detecting workflow-level anomalies in modern agentic AI systems. Chapter 2 tackles data quality heterogeneity in FL to improve global model utility. Noisy-labeled and imbalanced local data among clients can severely hinder training efficiency and model convergence. We propose a quality and fairness-aware client selection mechanism based on a novel Quality-of-Model (QoM) metric that evaluates client contribution without requiring access to model updates or gradients. Our approach prioritizes clients with high-quality contributions while ensuring diversity through a sortition-inspired randomized selection process, improving convergence speed and reducing communication overhead. Chapter 3 addresses Byzantine resilience to ensure FL trustworthiness from a system security perspective. FL's distributed nature forces the central server to blindly trust local training processes, making it vulnerable to model poisoning and data poisoning attacks from malicious participants. We propose a remote attestation-based approach that regains transparency into client-side training by verifying computation integrity via Trusted Execution Environment (TEE)-generated cryptographic attestation reports, enabling the server to reject malicious updates while preserving performance under non-IID data distributions. Building on these foundations, Chapter 4 demonstrates FL applied to network intrusion detection. We propose a geometric feature learning approach that projects network traffic into a compact, well-structured representation space, combining contrastive feature learning with H-Score optimization to maximize intra-class compactness and inter-class separability. The resulting federated intrusion detection system supports both anomaly detection and precise attack type identification, including zero-day threat exploration through entropy-based uncertainty analysis. Chapter 5 addresses a distinct and critical privacy threat in FL: model inversion attacks (MIAs) that allow a malicious server to reconstruct private training data from shared model updates. We identify that the reconstruction success of all known MIAs is fundamentally bounded by the local batch size relative to the model's first-layer leakage capacity. Building on this observation, we propose Aegis, a defense that synthesizes an auxiliary dataset to push the effective batch size beyond this capacity, collapsing the server's closed-form reconstructions without modifying the FL protocol or perturbing patient data. We demonstrate the approach on medical imaging datasets. Chapter 6 extends our security focus beyond FL to the rapidly growing domain of agentic AI. Modern agentic systems execute complex tasks through long-horizon workflows involving multi-agent coordination and tool invocation, creating a new risk surface where a single injected or erroneous step propagates through downstream dependencies. We present Skynet, a workflow-level anomaly detection framework that models multi-agent execution as directed workflow graphs and learns benign behavior jointly over semantic and structural dimensions. Trained exclusively on benign workflows, Skynet detects both adversarial manipulations and intrinsic execution failures under a single decision rule, naturally extending to zero-day anomalies. Together, these five chapters provide a comprehensive framework for building trustworthy machine learning systems: improving FL utility and trustworthiness, applying FL to network and medical security, and detecting anomalies in emerging agentic AI architectures, advancing the state of the art in machine learning security in adversarial and distributed environments. | en |
| dc.description.abstractgeneral | Machine learning (ML) now powers decisions that touch our daily lives, from detecting diseases in hospitals to defending computer networks from cyberattacks. But as these systems handle more sensitive data and operate across more organizations, a pressing question arises: can we actually trust them to be safe, secure, and private? This dissertation answers that question by designing trustworthy ML systems for real-world security applications, centered on two fast-growing technologies: federated learning and agentic AI. Federated learning lets many devices or organizations train a shared AI model together without ever revealing their private data. Imagine dozens of hospitals collaborating to build a better cancer-detection model while each hospital's patient records never leave its walls. This promise is powerful, but it also opens new doors for attackers: a dishonest participant can secretly poison the shared model, and a curious server can try to steal private data hidden inside the updates it receives. The chapters of this dissertation close these doors one by one and then look ahead to the next wave of AI risks. Chapter 2 confronts a practical reality: in the real world, data is messy, mislabeled, and unevenly distributed. We build a smart participant-selection method that automatically identifies and rewards the most helpful contributors, so the shared model learns faster and wastes less network bandwidth, making collaborative AI cheaper and more practical to deploy at scale. Chapter 3 tackles sabotage. Because the coordinating server normally cannot see what happens on each participant's device, a single bad actor can quietly corrupt the model that everyone relies on. We use tamper-proof hardware to verify that each participant trained honestly before its work is accepted, giving organizations the confidence to collaborate even with partners they do not fully trust. Chapter 4 brings these ideas to cybersecurity itself. We develop an intrusion detection system that learns to tell normal network traffic apart from malicious activity and can even raise the alarm on brand-new kinds of cyberattacks it has never encountered, helping defenders stay ahead of constantly evolving threats. Chapter 5 protects patient privacy in medical AI. We show that an adversarial server can reconstruct recognizable patient images from the model updates hospitals share, a serious privacy breach, and then neutralize the attack by masking the real training signal with synthetic data, all without disrupting the learning process or slowing diagnosis. Chapter 6 looks to the future of AI safety. Increasingly, teams of AI agents work together over long chains of steps to complete complex tasks, and a single corrupted or mistaken step can cascade into large-scale failures. We build a monitoring framework that learns what healthy multi-agent behavior looks like and instantly flags anything abnormal, catching both deliberate attacks and honest mistakes before they spread. Together, these chapters advance the science and practice of building AI systems that are robust to messy data, resilient against attackers, protective of personal privacy, and capable of watching over their own behavior, helping bring trustworthy AI from the lab into the high-stakes settings where people depend on it. | en |
| dc.description.degree | Doctor of Philosophy | en |
| dc.format.medium | ETD | en |
| dc.identifier.other | vt_gsexam:47502 | en |
| dc.identifier.uri | https://hdl.handle.net/10919/143700 | en |
| dc.language.iso | en | en |
| dc.publisher | Virginia Tech | en |
| dc.rights | In Copyright | en |
| dc.rights.uri | http://rightsstatements.org/vocab/InC/1.0/ | en |
| dc.subject | Machine Learning Security | en |
| dc.subject | Federated Learning | en |
| dc.subject | Privacy-Preserving Machine Learning | en |
| dc.subject | Agentic AI | en |
| dc.subject | Anomaly Detection | en |
| dc.title | Building Trustworthy Machine Learning Systems for Security: From Federated Learning to Agentic AI | en |
| dc.type | Dissertation | en |
| thesis.degree.discipline | Computer Science & Applications | en |
| thesis.degree.grantor | Virginia Polytechnic Institute and State University | en |
| thesis.degree.level | doctoral | en |
| thesis.degree.name | Doctor of Philosophy | en |