Expert-Driven Evaluation Platforms for Scientific Information Extraction Using Large Language Models
Files
TR Number
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Researchers are increasingly adopting Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) to extract scientific information in structured formats for training ML models. However, generative models are stochastic and prone to hallucinations. This necessitates expert-driven evaluation to verify factual accuracy and scientific reasoning. Traditional evaluation workflows rely on generic survey tools and spreadsheets. These tools impose cognitive load and may introduce psychological biases into the evaluation process. This thesis presents the design, architecture, and deployment of two visual analytics platforms to address these cognitive and operational challenges in LLM evaluation. Specifically, these platforms support the evaluation of extraction pipelines developed for computational virology. These underlying computational models are designed to parse academic literature to identify mutations enabling zoonotic host shifts and to extract quantitative viral inactivation metrics. The VILLA Evaluation Platform provides a structured environment for the qualitative assessment of LLM-generated reasoning regarding viral mutations. By implementing interface strategies such as blind evaluation, randomized presentation, and side-by-side component layouts, VILLA mitigates anchoring and central tendency biases. The Persist Extraction Interface serves as a diagnostic tool for evaluating the extraction of structured, quantitative data. The Persist interface aligns model outputs directly with human-verified ground-truth datasets to capture feedback across different models and pipeline iterations. These process-oriented interfaces translate raw probabilistic outputs into the actionable diagnostic data required by machine learning engineers. By standardizing collaborative workflows, this work provides a scalable foundation for assessing and optimizing research agents in specialized scientific domains.