Risk-averse learning with non-stationary distributionsShow others and affiliations
2026 (English)In: Automatica, ISSN 0005-1098, E-ISSN 1873-2836, Vol. 190, article id 113060Article in journal (Refereed) Published
Abstract [en]
Considering non-stationary environments in online optimization enables decision-makers to effectively adapt to changes and improve their performance over time. In such cases, it is favorable to adopt a strategy that minimizes the negative impact of change to avoid potentially risky situations. In this paper, we investigate risk-averse online optimization where the distribution of random costs changes over time. The Conditional Value at Risk (CVaR) is employed as risk measure. Due to the difficulty of obtaining the exact CVaR gradient, we employ a zeroth-order approach that queries the cost values multiple times per iteration and estimates the CVaR gradient from these samples. In regret analysis, the varying distributions are captured by a novel variation metric based on the Wasserstein distance. Given that the distribution variation is sublinear in the iteration horizon, we show that the developed learning algorithm achieves sublinear dynamic regret with high probability for both convex and strongly convex functions. Moreover, theoretical results suggest that dynamic regret bounds decrease with increasing sampling numbers until they reach a specific limit. Finally, we provide numerical experiments of dynamic pricing in a parking lot to illustrate the efficacy of the designed algorithm.
Place, publisher, year, edition, pages
Elsevier BV , 2026. Vol. 190, article id 113060
Keywords [en]
Dynamic regret, Online convex optimization, Risk-averse, Time-varying distribution
National Category
Computer Sciences Probability Theory and Statistics
Identifiers
URN: urn:nbn:se:kth:diva-382822DOI: 10.1016/j.automatica.2026.113060Scopus ID: 2-s2.0-105038838361OAI: oai:DiVA.org:kth-382822DiVA, id: diva2:2064538
Note
QC 20260602
2026-06-022026-06-022026-06-02Bibliographically approved