Multi-Stage Modeling with Gaussian Processes

dc.contributor.authorFlowers, Anna Rabrenen
dc.contributor.committeechairGramacy, Robert B.en
dc.contributor.committeechairFranck, Christopher Thomasen
dc.contributor.committeememberKrometis, Justin Augusten
dc.contributor.committeememberDeHart, Stephanie Pickleen
dc.contributor.departmentStatisticsen
dc.date.accessioned2026-07-10T08:00:43Zen
dc.date.available2026-07-10T08:00:43Zen
dc.date.issued2026-07-09en
dc.description.abstractGaussian processes (GPs) furnish accurate nonlinear predictions with well-calibrated uncertainty. However, GPs alone are not always flexible enough to model complex real-world phenomena. To remedy this issue, it is beneficial to chain together multiple models. These so-called multi-stage models fit models in sequence, using output from one model as training data in the next model. I utilize multi-stage modeling with GPs to achieve two tasks. First, I improve GP prediction accuracy for data from processes with sudden changes, or "jumps,"' in the output variable by creating a new cluster-based (latent) feature and adding it to the input matrix. Then, I fit a GP to understand the circumstances under which a machine learning (ML) model performs best, and use that GP to propose the best circumstances in which to make future data acquisitions. I do this by treating the composition of metadata and the performance of the ML model, respectively, as the inputs and output of a GP. I vet both methods on a selection of real and synthetic benchmark examples from the recent literature.en
dc.description.abstractgeneralGaussian processes (GPs) are statistical models chosen for their accurate predictions and estimates of uncertainty. However, GPs alone are not always flexible enough to model complex real-world phenomena. To remedy this issue, it is beneficial to combine GPs with other types of models. These so-called multi-stage models fit models one after the other, incorporating information from one model to create the next. I utilize multi-stage modeling with GPs to achieve two tasks. First, I improve GP predictions for data from processes with sudden changes, or "jumps," in the response by using an initial model to learn where those jumps occur. Then, I use a GP to understand the types of data that yield the best machine learning model performance. This can be used to guide future data acquisitions and ensure the resulting model is as good as possible. I do this by relating the balance of the metadata to the performance of the machine learning model. I vet both tasks on a collection of real and synthetic examples.en
dc.description.degreeDoctor of Philosophyen
dc.format.mediumETDen
dc.identifier.othervt_gsexam:46395en
dc.identifier.urihttps://hdl.handle.net/10919/143624en
dc.language.isoenen
dc.publisherVirginia Techen
dc.rightsCreative Commons Attribution 4.0 Internationalen
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/en
dc.subjectcomputer experimenten
dc.subjectnonstationarityen
dc.subjectactive learningen
dc.subjectmetadataen
dc.titleMulti-Stage Modeling with Gaussian Processesen
dc.typeDissertationen
thesis.degree.disciplineStatisticsen
thesis.degree.grantorVirginia Polytechnic Institute and State Universityen
thesis.degree.leveldoctoralen
thesis.degree.nameDoctor of Philosophyen

Files

Original bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Flowers_AR_D_2026.pdf
Size:
3.6 MB
Format:
Adobe Portable Document Format