kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Risk-averse learning with non-stationary distributions
KTH, School of Electrical Engineering and Computer Science (EECS), Decision and Control Systems.ORCID iD: 0000-0003-1146-2473
KTH, School of Electrical Engineering and Computer Science (EECS), Decision and Control Systems.ORCID iD: 0000-0001-6464-492X
College of Electronics and Information Engineering, State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University, Shanghai, 200092, China.
Mechanical Engineering and Material Science, Duke University, Durham, NC 27708, USA.
Show others and affiliations
2026 (English)In: Automatica, ISSN 0005-1098, E-ISSN 1873-2836, Vol. 190, article id 113060Article in journal (Refereed) Published
Abstract [en]

Considering non-stationary environments in online optimization enables decision-makers to effectively adapt to changes and improve their performance over time. In such cases, it is favorable to adopt a strategy that minimizes the negative impact of change to avoid potentially risky situations. In this paper, we investigate risk-averse online optimization where the distribution of random costs changes over time. The Conditional Value at Risk (CVaR) is employed as risk measure. Due to the difficulty of obtaining the exact CVaR gradient, we employ a zeroth-order approach that queries the cost values multiple times per iteration and estimates the CVaR gradient from these samples. In regret analysis, the varying distributions are captured by a novel variation metric based on the Wasserstein distance. Given that the distribution variation is sublinear in the iteration horizon, we show that the developed learning algorithm achieves sublinear dynamic regret with high probability for both convex and strongly convex functions. Moreover, theoretical results suggest that dynamic regret bounds decrease with increasing sampling numbers until they reach a specific limit. Finally, we provide numerical experiments of dynamic pricing in a parking lot to illustrate the efficacy of the designed algorithm.

Place, publisher, year, edition, pages
Elsevier BV , 2026. Vol. 190, article id 113060
Keywords [en]
Dynamic regret, Online convex optimization, Risk-averse, Time-varying distribution
National Category
Computer Sciences Probability Theory and Statistics
Identifiers
URN: urn:nbn:se:kth:diva-382822DOI: 10.1016/j.automatica.2026.113060Scopus ID: 2-s2.0-105038838361OAI: oai:DiVA.org:kth-382822DiVA, id: diva2:2064538
Note

QC 20260602

Available from: 2026-06-02 Created: 2026-06-02 Last updated: 2026-06-02Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Wang, SiyiWang, ZifanJohansson, Karl Henrik

Search in DiVA

By author/editor
Wang, SiyiWang, ZifanJohansson, Karl Henrik
By organisation
Decision and Control Systems
In the same journal
Automatica
Computer SciencesProbability Theory and Statistics

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 9 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf