Explainable Machine Learning Framework for Accurate Reference-free Biological Data Deconvolution
Files
TR Number
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Bulk omics data captures mixed molecular signals from multicellular tissue compositions, posing significant challenges in accurate interpretation of biological changes over samples. Reference free deconvolution offers a flexible framework to estimate cell type proportions and specific expressions from bulk data, thus uncovering latent cellular and molecular architecture of tissue ecosystems without relying on predefined references. However, existing deconvolution methods are highly sensitive to several hidden confounders, including asymmetric gene expression, informative missingness, inter-cell-type imbalance, and deviation from identifiability conditions. These issues also propagate across preprocessing and modeling stages, collectively, leading to reduced accuracy of bulk deconvolution and downstream inference. In this dissertation, we present a comprehensive methodological framework with effective workflow for reference-free deconvolution of complex biological data. The proposed workflow integrates four key components spanning preprocessing, missing value imputation, structural correction, and discriminative deconvolution. First, we propose Cosbin, an iterative normalization strategy that identifies consistently expressed genes and removes asymmetrically differentially expressed genes, thereby preserving the geometric structure of the data. Second, we propose mechanism-integrated group-wise pre-imputation, which explicitly models multiple missingness mechanisms and preserves biologically informative missing patterns, particularly for marker-like genes. Third, we introduce iterative equilibration of cell-type expression profiles to correct inter cell-type asymmetry and improve the identifiability and accuracy of proportion estimation. Fourth, we apply and evaluate CAM3.0, an enhanced convex geometry-based unsupervised deconvolution algorithm, to estimate latent molecular archetypes and compositions on diverse omics data types from real biological bulk samples. To further improve deconvolution accuracy and efficiency particularly when constituent cell types are highly mixed or hardly separable, we also propose and develop an effective cosine similarity based discriminative analysis of mixtures method (csDAM), specifically to deconvolute highly mixed bulk expression data. Facilitated by the monotonic relationship between signature gene specificity and cosine similarity rank distribution, csDAM achieves highly accurate estimation of cell type proportions and specific expressions. Simulation and real-data based studies show both improved deconvolution performance and computational efficiency by csDAM compared to most relevant peer methods. In summary, the work in this dissertation addresses several major limitations of existing reference free deconvolution approaches by collectively integrating improved preprocessing, missing value imputation, structural correction, and discriminative modeling into a unified framework. Extensive simulations and real-data applications demonstrate improved performance in terms of stability, accuracy, and interpretability under challenging conditions, including high noise, complex missingness, and strong cellular heterogeneity. This study highlights the critical role of structural considerations in deconvolution and provides a scalable solution for extracting biologically meaningful latent features and archetypes from large-scale bulk omics data.