Towards Understanding the Role of Feedback in Online Learning

Loading...
Thumbnail Image

TR Number

Date

2026-07-02

Journal Title

Journal ISSN

Volume Title

Publisher

Virginia Tech

Abstract

Online learning studies sequential decision-making under uncertainty, where a learner repeatedly interacts with an environment and updates their decisions based on observed feedback. Across a wide range of applications—including recommendation systems, online advertising, reinforcement learning, and large-scale networked infrastructures—the nature of feedback fundamentally determines what information is revealed, what guarantees are achievable, and how efficiently learning can proceed. Despite its central importance, existing theory typically assumes a fixed feedback model, resulting in fragmented analyses tailored to specific settings. This dissertation develops a unified, feedback-centric perspective on online learning, systematically investigating how distinct feedback regimes shape regret minimization, robustness, and algorithm design.

We first study how the amount of feedback influences learnability in online learning with switching costs. Moving beyond the classical bandit versus full-information dichotomy, we consider intermediate regimes in which additional observations are available under explicit budget constraints. By characterizing the minimax regret as a function of the feedback budget, we establish a sharp phase-transition phenomenon: when feedback is scarce, regret matches the bandit-optimal rate, whereas beyond a critical threshold, additional feedback provably improves achievable performance. These results provide a quantitative understanding of how feedback availability reshapes the statistical complexity of online learning with switching costs.

We next investigate environments where feedback exhibits heavy-tailed noise and unbounded variability. In such settings, the challenge lies not in the quantity of feedback but in its statistical properties. We design learning algorithms that remain robust under heavy-tailed losses while simultaneously adapting between stochastic and adversarial regimes. Our analysis shows that the tail heaviness and noise level of feedback fundamentally determine achievable regret guarantees, and that carefully constructed algorithms can attain best-of-both-worlds performance even when feedback is highly irregular.

We then analyze hybrid feedback models in reinforcement learning, where learners simultaneously receive on-policy and off-policy information in adversarial Markov decision processes. We develop a unified framework that integrates these heterogeneous feedback sources, achieving coverage-dependent improvements when auxiliary data are informative while preserving worst-case guarantees. This work highlights how the source of feedback influences efficiency in sequential decision-making.

Finally, we extend this investigation to structured and decentralized feedback in large-scale networked systems through an online learning formulation of end-to-end (E2E) service level agreement (SLA) decomposition for 5G/6G network slicing. In this setting, the learner observes domain-level performance feedback while the underlying system dynamics remain unknown and distributed across heterogeneous domains, and the E2E performance is determined by some composition of domain-level ones, which essentially exhibits a hierarchical nature. We develop a provably efficient learning framework and validate it through trace-driven and testbed-based evaluations, demonstrating how feedback granularity and system-level structure affect convergence, regret, and constraint satisfaction in practice.

Together, these results establish general principles for learning under diverse feedback regimes (varying in amount, statistical properties, source, and structure) bridging theoretical foundations with practical learning systems and advancing a unified understanding of the role of feedback in online learning.

Description

Keywords

Online learning, regret minimization, multi-armed bandits, reinforcement learning

Citation