Data-Driven Prokaryotic Genome Identification: LIN Assignment, Taxonomy Correspondence, and Deployment

Loading...
Thumbnail Image

TR Number

Date

2026-08-21

Journal Title

Journal ISSN

Volume Title

Publisher

Virginia Tech

Abstract

This dissertation addresses a central challenge in microbial genomics: how to organize and identify rapidly growing genome collections in a way that is both computationally scalable and biologically interpretable. It presents a data-driven framework centered on Life Identification Numbers (LINs), where each genome receives a hierarchical, threshold-based label, and shared prefixes encode relatedness at multiple resolutions. The work develops LINflow 2.0, a modular assignment system that combines high-throughput sketch-based filtering with average nucleotide identity refinement to support efficient, incremental LIN assignment at database scale. LIN-derived clusters are then compared against NCBI and GTDB taxonomies, showing strong correspondence at many ranks, while also revealing lineage-specific discordances that can guide recircumscription. The framework is operationalized in genomeRxiv, a web platform that supports rapid genome or sketch submission, query assignment, taxonomy circumscriptions, and search across related genomes. Together, these theoretical, algorithmic, and platform contributions provide a unified approach that bridges traditional taxonomy and strain-level typing, enabling faster and more reproducible genome identification for microbial research, surveillance, and outbreak response.

Description

Keywords

Bacteria, Archaea, taxonomy, genomics, k-mers, average nucleotide identity, Jaccard similarity, LIN robustness, sourmash

Citation