kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
On the effect of clock offsets and quantization on learning-based adversarial games
School of Aerospace Engineering, Georgia Institute of Technology, Atlanta, GA, USA.
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Information Science and Engineering.ORCID iD: 0000-0001-5983-0875
School of Aerospace Engineering, Georgia Institute of Technology, Atlanta, GA, USA.
Carnegie Mellon University/Software Engineering Institute, Pittsburgh, PA, USA.
2024 (English)In: Automatica, ISSN 0005-1098, E-ISSN 1873-2836, Vol. 167, article id 111762Article in journal (Refereed) Published
Abstract [en]

In this work, we consider systems whose components suffer from clock offsets and quantization and study the effect of those on a reinforcement learning (RL) algorithm. Specifically, we consider an off-policy iterative RL algorithm for continuous-time systems, which uses input and state data to approximate the Nash-equilibrium of a zero-sum game. However, the data used by this algorithm are not consistent with one another, in that each of them originates from a slightly different time instant of the past, hence putting the convergence of the algorithm in question. We prove that, given that these timing inconsistencies remain below a certain threshold, the iterative off-policy RL algorithm will still converge epsilon-closely to the desired Nash policy. However, this result is conditional to a certain Lipschitz continuity and differentiability condition on the input-state data collected, which is indispensable in the presence of clock offsets. A similar result is also derived when quantization of the measured state is considered. Finally, unlike prior work, we provide a sufficiently rich data condition for the execution of the iterative RL algorithm, which can be verified a priori across all iteration indices. Simulations are performed, which verify and clarify theoretical findings.

Place, publisher, year, edition, pages
Elsevier BV , 2024. Vol. 167, article id 111762
Keywords [en]
Clock offsets, Learning, Quantization, Zero-sum games
National Category
Control Engineering
Identifiers
URN: urn:nbn:se:kth:diva-348313DOI: 10.1016/j.automatica.2024.111762ISI: 001257972700001Scopus ID: 2-s2.0-85195608324OAI: oai:DiVA.org:kth-348313DiVA, id: diva2:1874685
Note

QC 20240624

Available from: 2024-06-20 Created: 2024-06-20 Last updated: 2024-07-15Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Kanellopoulos, Aris

Search in DiVA

By author/editor
Kanellopoulos, Aris
By organisation
Information Science and Engineering
In the same journal
Automatica
Control Engineering

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 99 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf