kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Policy Evaluation in Distributional LQR
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Decision and Control Systems (Automatic Control). KTH, School of Electrical Engineering and Computer Science (EECS), Centres, Digital futures.ORCID iD: 0000-0001-6464-492X
Imperial College London, Department of Electrical and Electronic Engineering, London, U.K.ORCID iD: 0000-0003-2338-5487
Technical University of Munich, School of Computation, Information and Technology, Munich, Germany.ORCID iD: 0000-0003-1146-2473
Duke University, Department of Mechanical Engineering and Materials Science, Durham, NC, USA.ORCID iD: 0000-0003-1748-8228
Show others and affiliations
2025 (English)In: IEEE Transactions on Automatic Control, ISSN 0018-9286, E-ISSN 1558-2523, Vol. 70, no 11, p. 7477-7492Article in journal (Refereed) Published
Abstract [en]

Distributional reinforcement learning (DRL) enhances the understanding of the effects of the randomness in the environment by letting agents learn the distribution of a random return, rather than its expected value as in standard reinforcement learning. Meanwhile, a challenge in DRL is that the policy evaluation typically relies on the representation of the return distribution, which needs to be carefully designed. In this paper, we address this challenge for the special class of DRL problems that rely on a discounted linear quadratic regulator (LQR), which we call distributional LQR. Specifically, we provide a closed-form expression for the distribution of the random return, which is applicable for all types of exogenous disturbance as long as it is independent and identically distributed (i.i.d.). We show that the variance of the random return is bounded if the fourth moment of the exogenous disturbance is bounded. Furthermore, we investigate the sensitivity of the return distribution to model perturbations. While the proposed exact return distribution consists of infinitely many random variables, we show that this distribution can be well approximated by a finite number of random variables. The associated approximation error can be analytically bounded under mild assumptions. When the model is unknown, we propose a model-free approach for estimating the return distribution, supported by sample complexity guarantees. Finally, we extend our approach to partially observable linear systems. Numerical experiments are provided to illustrate the theoretical results.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE) , 2025. Vol. 70, no 11, p. 7477-7492
Keywords [en]
distribution sensitivity, Distributional LQR, distributional RL, partially observable system, policy evaluation
National Category
Probability Theory and Statistics
Identifiers
URN: urn:nbn:se:kth:diva-366183DOI: 10.1109/TAC.2025.3575649ISI: 001605046600002Scopus ID: 2-s2.0-105007415091OAI: oai:DiVA.org:kth-366183DiVA, id: diva2:1981919
Note

QC 20260127

Available from: 2025-07-07 Created: 2025-07-07 Last updated: 2026-01-27Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Wang, ZifanJohansson, Karl H.

Search in DiVA

By author/editor
Wang, ZifanGao, YulongWang, SiyiZavlanos, Michael M.Abate, AlessandroJohansson, Karl H.
By organisation
Decision and Control Systems (Automatic Control)Digital futures
In the same journal
IEEE Transactions on Automatic Control
Probability Theory and Statistics

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 52 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf