LLM-Guided Run-Time Parameter Optimization for Energy-Efficient Model Inference
| dc.contributor.author | Crumpacker, Mary Katelyn | en |
| dc.contributor.committeechair | Nikolopoulos, Dimitrios S. | en |
| dc.contributor.committeemember | Ellis, Margaret O.'Neil | en |
| dc.contributor.committeemember | Back, Godmar Volker | en |
| dc.contributor.department | Computer Science and#38; Applications | en |
| dc.date.accessioned | 2026-07-09T08:00:16Z | en |
| dc.date.available | 2026-07-09T08:00:16Z | en |
| dc.date.issued | 2026-07-08 | en |
| dc.description.abstract | The scale of Large Language Models (LLMs) has been increasing over time, and is predicted to continue to increase as LLMs become an integral part of many real world workflows. However, LLMs consume a tremendous amount of energy, which becomes a large concern in the scale of the demand for these tools. Inference engines and serving systems, like vLLM and PyTorch, are used to perform efficient LLM inference. These tools include runtime parameters that have the potential to improve the energy efficiency of inference. The task is choosing the values for those parameters that actually result in lower energy consumption. However, these parameters have complex relationships and trade offs that make this choice difficult and sometimes non intuitive. This choice can also require deep knowledge of the application or inefficient traditional optimization methods. In this work, we created a human-in-the-loop flow with LLM assisted runtime parameter optimization in order to solve this issue. With human created, specific feedback prompting methods, chat based LLMs can iteratively find energy efficient inference parameters faster than traditional search methods. We evaluated LLM as an optimization method on a text prompt workload with vLLM as the inference engine and on an image prompt workload with PyTorch as the inference engine. Across these two systems, we compared the performance of LLM optimization when prompting with baseline prompts that only gave the minimum amount of information and enhanced prompts which included different prompting strategies with the goal of improving performance. LLM optimization that uses the enhnaced prompt strategy, was compared to a more traditional optimization method: Sobol sampling. The enhanced prompting strategy outperforms the baseline across all experiment runs. It achieves an average of 45% reduction in energy-per-token for vLLM inference and an average of 90.7% reduction for PyTorch multimodal inference relative to the default configurations. Compared to Sobol sampling, the LLM guided approach reaches lower energy-per-token values more efficiently in the vLLM inference and achives lower energy-per-image values in the PyTorch setting. Sobol sampling also consistently failed to avoid out of memory configurations in the PyTorch configuration. These results suggest that LLMs may serve as effective optimization agents for reducing the energy consumption of AI inference, without requiring expertise or exhaustive search. | en |
| dc.description.abstractgeneral | Large Language Models (LLMs), which are AI systems that generate human-like text, have become an integral part of many real world workflows. However, LLMs consume a lot of energy, which becomes a large concern because of their high demand. As these models become integrated into different workflows, applications have been developed to handle the growing number of users that are sending prompts to the model. These are called inference engines. This introduces the challenge of selecting inference engine settings to reduce energy consumption. Oftentimes this requires deep knowledge of the application or methods that systematically try different settings to find the best one, which can take days. In this work, we created a workflow where humans guide the model as it selects settings, in order to solve this issue. With iterative, human written prompts, LLMs can find energy efficient settings faster than the systematic methods. LLMs also have the potential to tailor their solutions to different applications and hardware setups while also keeping in mind any user provided constraints. We evaluate performance using both text based and image based tasks. For both approaches, we used simple human written prompts, called baseline prompts, and more complicated human written prompts, called enhanced prompts, to guide the LLM in finding the energy efficient settings. Our results show that more detailed human guidance helps the system find energy efficient settings faster than both simpler guidance and traditional systematic methods. This suggests that human AI collaboration can be an effective way to reduce the energy cost of running large AI models while still maintaining performance. | en |
| dc.description.degree | Master of Science | en |
| dc.format.medium | ETD | en |
| dc.identifier.other | vt_gsexam:46961 | en |
| dc.identifier.uri | https://hdl.handle.net/10919/143610 | en |
| dc.language.iso | en | en |
| dc.publisher | Virginia Tech | en |
| dc.rights | In Copyright | en |
| dc.rights.uri | http://rightsstatements.org/vocab/InC/1.0/ | en |
| dc.subject | Large language models | en |
| dc.subject | energy-efficient inference | en |
| dc.subject | human-in-the-loop optimization | en |
| dc.subject | hardware-aware tuning | en |
| dc.subject | runtime parameter optimization | en |
| dc.title | LLM-Guided Run-Time Parameter Optimization for Energy-Efficient Model Inference | en |
| dc.type | Thesis | en |
| thesis.degree.discipline | Computer Science & Applications | en |
| thesis.degree.grantor | Virginia Polytechnic Institute and State University | en |
| thesis.degree.level | masters | en |
| thesis.degree.name | Master of Science | en |