Toward Trustworthy Health AI Systems: Advancing Clinician–AI Interaction Through Interpretability, Fairness, and Context-Aware Multi-Agent Safety Architectures

Loading...
Thumbnail Image

TR Number

Date

2026-08-31

Journal Title

Journal ISSN

Volume Title

Publisher

Virginia Tech

Abstract

AI has become increasingly integrated into healthcare, supporting clinical decision making, patient risk prediction, and patient-facing information systems. Despite substantial advances in predictive modeling and LLMs, widespread adoption of AI in healthcare remains constrained by challenges related to interpretability, fairness, safety, and human- AI interaction. Healthcare applications require not only accurate predictions and recommendations but also transparent, equitable, and clinically reliable systems that can be trusted by end-users. This dissertation advances the design of trustworthy health AI systems through three complementary studies focused on clinician–AI interaction, fairness-aware clinical risk prediction, and context-aware multi-agent safety architectures. In essay 1, a systematic review of explainable AI (XAI) in healthcare synthesizes existing approaches and proposes a clinician-centered framework for improving collaboration between AI systems and healthcare professionals. The review identifies key challenges in implementing interpretable AI within clinical decision support systems and provides a roadmap for responsible deployment. In the second essay, a fair and clinically interpretable machine learning framework is developed to predict distinct opioid-related respiratory deterioration events among hospitalized patients receiving opioid therapy. Using electronic health record (HER) data, the proposed multiclass framework differentiates Naloxone intervention events, Blue Code respiratory arrest events, and Rapid Response events. The framework integrates explainability methods, fairness evaluation, and age-specific threshold optimization to improve detection of severe respiratory outcomes while maintaining clinical interpretability. Experimental results demonstrate substantial performance improvements compared with traditional clinical risk scoring approaches. The third essay, a context-aware multi-agent safety architecture, CareGuardAI, is introduced to address clinical safety risks and hallucination risks in patient-facing medical LLMs. The proposed system combines safety-constrained generation, risk assessment agents, and iterative refinement mechanisms to evaluate both medical safety and factual reliability before responses are delivered to patients. Across multiple healthcare safety and hallucination benchmarks, the architecture demonstrates improved performance relative to strong baseline models while maintaining bounded response latency. Collectively, these studies contribute a unified perspective on trustworthy health AI by integrating interpretability, fairness, and safety into the design of healthcare AI systems. The findings provide practical and methodological guidance for developing AI-enabled clinical decision support systems that improve effective human–AI interaction (clinicians and patients) and responsible deployment in high-stakes healthcare environments.

Description

Keywords

Trustworthy AI, Health AI Systems, Clinician–AI Interaction, Clinical Decision Support Systems, Interpretable Machine Learning, Fairness-Aware AI, Context-Aware AI, Multi- Agent Systems, Large Language Models, AI Safety, Responsible AI, Agentic AI

Citation