Understanding and Enhancing Sequential Decision-Making and Alignment in Foundation Models
| dc.contributor.author | Sel, Bilgehan | en |
| dc.contributor.committeechair | Jin, Ming | en |
| dc.contributor.committeemember | Ramakrishnan, Narendran | en |
| dc.contributor.committeemember | Jia, Ruoxi | en |
| dc.contributor.committeemember | Yanardag Delul, Pinar | en |
| dc.contributor.committeemember | Soysal, Alkan | en |
| dc.contributor.department | Electrical and Computer Engineering | en |
| dc.date.accessioned | 2026-08-27T08:00:32Z | en |
| dc.date.available | 2026-08-27T08:00:32Z | en |
| dc.date.issued | 2026-08-26 | en |
| dc.description.abstract | Foundation models are used as sequential decision-makers, in tasks whose outputs unfold as dependent steps: reasoning, planning, decisions affecting several parties, and generation under safety requirements. This dissertation asks which factors govern the performance of foundation models in such settings and how that performance can be improved. The answer locates those factors in training data and training procedure. Pretraining text preserves finished solutions far more often than the failed attempts behind them, so a model trained on it defaults to confident, linearly coherent continuation; exploring alternative paths and retracting mistaken steps are underrepresented behaviors, not absent capabilities. The contributions elicit these behaviors in context and then teach them by supervised and reinforcement learning. One observation recurs: a model given room to explore diverse solution paths and to back out of its own mistakes makes better decisions, and in safety-critical settings the same capacity lets it recover from unsafe trajectories. The body has three parts. Part I treats foundation models in reinforcement learning: a meta-learning method for sequences of derivative-free optimization tasks, with task-averaged regret guarantees; policy optimization under several reward objectives and hard safety constraints, with a rectification step that restores feasibility after a detected violation; and an analysis tracing in-context reinforcement learning to the diversity of the pretraining task distribution. Part II treats large language models in sequential decision-making: a tool-use framework that translates natural-language energy-management requests into solver-ready optimization programs; the Algorithm of Thoughts, a prompting strategy whose exemplars record a search process so that the model explores, prunes, and backtracks within a single generation; an extension to autonomous long-horizon planning; and a training pipeline that makes the behavior a concise default. Part III turns the same capacity to alignment and safety: a prompting framework that surveys a decision's consequences for every affected stakeholder before answering, and a reinforcement-learning method that trains the backtracking step as a safety signal, so that the model retracts an emerging violation and continues from the safe prefix. Recovery complements avoidance rather than replacing it; an integrated red-team study of a static classifier defense motivates judging safety on the generated trajectory. | en |
| dc.description.abstractgeneral | Modern artificial-intelligence systems are asked not only to answer single questions but to work through tasks that take many steps: solving a math problem, planning a sequence of actions, or writing a long answer. In tasks like these, the quality of the result depends on the whole chain of choices, not on any one of them. This dissertation asks what determines how well these systems handle such tasks and how they can be made to handle them better. Part of the answer lies in the text these systems learn from. Published writing mostly shows finished work: a textbook presents the proof that succeeded, not the attempts the author discarded along the way. A system trained on such text tends to keep writing forward with confidence, even after a wrong turn, because that is what its examples look like. Trying alternatives and going back to fix an earlier step are things these systems can do; their training gives them few examples of doing so. This dissertation shows that those habits can be supplied. Showing the system worked examples that include wrong turns and corrections, or training it on such behavior directly, improves how well it solves problems and makes plans. The same ability serves safety: a system that notices its response is turning harmful can take back the harmful part and continue safely, which works alongside teaching the system to avoid harm in the first place rather than replacing that teaching. The dissertation develops these methods, measures them against existing ones, and states where they apply and where they do not. | en |
| dc.description.degree | Doctor of Philosophy | en |
| dc.format.medium | ETD | en |
| dc.identifier.other | vt_gsexam:47535 | en |
| dc.identifier.uri | https://hdl.handle.net/10919/143769 | en |
| dc.language.iso | en | en |
| dc.publisher | Virginia Tech | en |
| dc.rights | In Copyright | en |
| dc.rights.uri | http://rightsstatements.org/vocab/InC/1.0/ | en |
| dc.subject | large language models | en |
| dc.subject | reinforcement learning | en |
| dc.subject | sequential-decision making | en |
| dc.subject | alignment | en |
| dc.subject | AI Safety | en |
| dc.title | Understanding and Enhancing Sequential Decision-Making and Alignment in Foundation Models | en |
| dc.type | Dissertation | en |
| thesis.degree.discipline | Computer Engineering | en |
| thesis.degree.grantor | Virginia Polytechnic Institute and State University | en |
| thesis.degree.level | doctoral | en |
| thesis.degree.name | Doctor of Philosophy | en |
Files
Original bundle
1 - 1 of 1