Multi-Stage Modeling with Gaussian Processes
| dc.contributor.author | Flowers, Anna Rabren | en |
| dc.contributor.committeechair | Gramacy, Robert B. | en |
| dc.contributor.committeechair | Franck, Christopher Thomas | en |
| dc.contributor.committeemember | Krometis, Justin August | en |
| dc.contributor.committeemember | DeHart, Stephanie Pickle | en |
| dc.contributor.department | Statistics | en |
| dc.date.accessioned | 2026-07-10T08:00:43Z | en |
| dc.date.available | 2026-07-10T08:00:43Z | en |
| dc.date.issued | 2026-07-09 | en |
| dc.description.abstract | Gaussian processes (GPs) furnish accurate nonlinear predictions with well-calibrated uncertainty. However, GPs alone are not always flexible enough to model complex real-world phenomena. To remedy this issue, it is beneficial to chain together multiple models. These so-called multi-stage models fit models in sequence, using output from one model as training data in the next model. I utilize multi-stage modeling with GPs to achieve two tasks. First, I improve GP prediction accuracy for data from processes with sudden changes, or "jumps,"' in the output variable by creating a new cluster-based (latent) feature and adding it to the input matrix. Then, I fit a GP to understand the circumstances under which a machine learning (ML) model performs best, and use that GP to propose the best circumstances in which to make future data acquisitions. I do this by treating the composition of metadata and the performance of the ML model, respectively, as the inputs and output of a GP. I vet both methods on a selection of real and synthetic benchmark examples from the recent literature. | en |
| dc.description.abstractgeneral | Gaussian processes (GPs) are statistical models chosen for their accurate predictions and estimates of uncertainty. However, GPs alone are not always flexible enough to model complex real-world phenomena. To remedy this issue, it is beneficial to combine GPs with other types of models. These so-called multi-stage models fit models one after the other, incorporating information from one model to create the next. I utilize multi-stage modeling with GPs to achieve two tasks. First, I improve GP predictions for data from processes with sudden changes, or "jumps," in the response by using an initial model to learn where those jumps occur. Then, I use a GP to understand the types of data that yield the best machine learning model performance. This can be used to guide future data acquisitions and ensure the resulting model is as good as possible. I do this by relating the balance of the metadata to the performance of the machine learning model. I vet both tasks on a collection of real and synthetic examples. | en |
| dc.description.degree | Doctor of Philosophy | en |
| dc.format.medium | ETD | en |
| dc.identifier.other | vt_gsexam:46395 | en |
| dc.identifier.uri | https://hdl.handle.net/10919/143624 | en |
| dc.language.iso | en | en |
| dc.publisher | Virginia Tech | en |
| dc.rights | Creative Commons Attribution 4.0 International | en |
| dc.rights.uri | http://creativecommons.org/licenses/by/4.0/ | en |
| dc.subject | computer experiment | en |
| dc.subject | nonstationarity | en |
| dc.subject | active learning | en |
| dc.subject | metadata | en |
| dc.title | Multi-Stage Modeling with Gaussian Processes | en |
| dc.type | Dissertation | en |
| thesis.degree.discipline | Statistics | en |
| thesis.degree.grantor | Virginia Polytechnic Institute and State University | en |
| thesis.degree.level | doctoral | en |
| thesis.degree.name | Doctor of Philosophy | en |
Files
Original bundle
1 - 1 of 1