Knowledge-Centric Multimodal Intelligence: Understanding, Reasoning, and Generation

Loading...
Thumbnail Image

Files

TR Number

Date

2026-06-30

Journal Title

Journal ISSN

Volume Title

Publisher

Virginia Tech

Abstract

Generative multimodal foundation models now handle a wide range of tasks spanning understanding, reasoning, and generation. Yet the knowledge they draw on at inference time is bounded by what their finite parameters memorized during training, leading to fabricated content, missed low-frequency details, and incomplete use of available evidence. This dissertation argues that closing the gap between what a model can access and what it actually uses drives capability gains across all three axes. For multimodal understanding, I develop a retrieval-augmented vision-language model that learns to ignore irrelevant retrievals, making knowledge-intensive visual question answering resilient to retriever noise. For reasoning, I first introduce a recursive questioning procedure that surfaces parametric knowledge a single-pass prompt cannot reach. I then address the complementary challenge that knowledge-driven methods alone cannot eliminate reasoning errors, developing self-revision mechanisms that let the model correct its own intermediate outputs in both autoregressive and discrete-diffusion generation paradigms. For multimodal generation, I extend retrieval augmentation into image synthesis, incorporating fine-grained visual evidence that the text prompt alone cannot specify. Across visual question answering, mathematical reasoning, constraint solving, and image generation, these methods outperform approaches that rely on parametric knowledge alone. The broader claim is that more effective utilization of available knowledge drives capability gains beyond what larger models or longer training alone can provide.

Description

Keywords

multimodal generative model, retrieval-augmented generation, reasoning

Citation