kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Online bandit non-cooperative games with arbitrary delays
Tongji University, The Department of Control Science and Engineering, China.
Tongji University, The Department of Control Science and Engineering, China, 201804; Tongji University, Shanghai Research Institute for Intelligent Autonomous Systems, the National Key Laboratory of Autonomous Intelligent Unmanned Systems & Frontiers Science Center for Intelligent Autonomous Systems, Ministry of Education, China.
Tongji University, The Department of Control Science and Engineering, China, 201804; Tongji University, Shanghai Research Institute for Intelligent Autonomous Systems, the National Key Laboratory of Autonomous Intelligent Unmanned Systems & Frontiers Science Center for Intelligent Autonomous Systems, Ministry of Education, China.
University of Toronto, The Department of Electrical and Computer Engineering, Toronto, Canada.
Show others and affiliations
2025 (English)In: 2025 IEEE 64th Conference on Decision and Control, CDC 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 6814-6819Conference paper, Published paper (Refereed)
Abstract [en]

This paper considers online bandit games with arbitrary delays, where the cost functions of all self-interested players are time-varying. In addition, players lack an explicit model of the game and can only learn their actions based on the sole available feedback of delayed cost values. To address this challenging setting, a novel learning algorithm named Cumulative Bandit Online Learning with arbitrary delays (CBOL-ad) is proposed. We conduct regret analysis for time-varying games where the player-specific problem is convex, explicitly revealing the influence of time delays and game structure on the regret bound. In particular, under certain delay conditions, our bound can achieve the same order as that of online bandit optimization problems without delays. Finally, numerical simulations are provided to illustrate the algorithmic performance.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE) , 2025. p. 6814-6819
National Category
Control Engineering Other Mathematics
Identifiers
URN: urn:nbn:se:kth:diva-378896DOI: 10.1109/CDC57313.2025.11312073Scopus ID: 2-s2.0-105031876553OAI: oai:DiVA.org:kth-378896DiVA, id: diva2:2051718
Conference
64th IEEE Conference on Decision and Control, CDC 2025, Rio de Janeiro, Brazil, Dec 9 2025 - Dec 12 2025
Note

Part of ISBN 9798331526276

QC 20260409

Available from: 2026-04-09 Created: 2026-04-09 Last updated: 2026-04-09Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Hu, Xiaoming

Search in DiVA

By author/editor
Hu, Xiaoming
By organisation
Numerical Analysis, Optimization and Systems Theory
Control EngineeringOther Mathematics

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 22 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf