kth.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Transfer-Entropy-Regularized Markov Decision Processes
Aerospace Engineering and Engineering Mechanics, University of Texas at Austin, Austin, TX, USA.
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Decision and Control Systems (Automatic Control).ORCID iD: 0000-0003-1835-2963
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Information Science and Engineering.ORCID iD: 0000-0002-7926-5081
2022 (English)In: IEEE Transactions on Automatic Control, ISSN 0018-9286, E-ISSN 1558-2523, Vol. 67, no 4, p. 1944-1951Article in journal (Refereed) Published
Abstract [en]

We consider the framework of transfer-entropy-regularized Markov Decision Process (TERMDP) in which the weighted sum of the classical state-dependent cost and the transfer entropy from the state random process to the control input process is minimized. Although TERMDPs are generally formulated as nonconvex optimization problems, an analytical necessary optimality condition can be expressed as a finite set of nonlinear equations, based on which an iterative forward-backward computational procedure similar to the Arimoto-Blahut algorithm is developed. It is shown that every limit point of the sequence generated by the proposed algorithm is a stationary point of the TERMDP. Applications of TERMDPs are discussed in the context of networked control systems theory and non-equilibrium thermodynamics. The proposed algorithm is applied to an information-constrained maze navigation problem, whereby we study how the price of information qualitatively alters the optimal decision polices. 

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE) , 2022. Vol. 67, no 4, p. 1944-1951
Keywords [en]
Entropy, Markov processes, Optimization, Random processes, Rate-distortion, Standards, Thermodynamics, Behavioral research, Computation theory, Iterative methods, Networked control systems, Nonlinear equations, Computational procedures, Markov Decision Processes, Navigation problem, Necessary optimality condition, Non equilibrium thermodynamics, Nonconvex optimization problem, Optimal decisions, Stationary points, Process control
National Category
Control Engineering
Identifiers
URN: urn:nbn:se:kth:diva-308832DOI: 10.1109/TAC.2021.3069347ISI: 000776167500027Scopus ID: 2-s2.0-85103769098OAI: oai:DiVA.org:kth-308832DiVA, id: diva2:1637930
Note

QC 20250512

Available from: 2022-02-15 Created: 2022-02-15 Last updated: 2025-05-12Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Sandberg, HenrikSkoglund, Mikael

Search in DiVA

By author/editor
Sandberg, HenrikSkoglund, Mikael
By organisation
Decision and Control Systems (Automatic Control)Information Science and Engineering
In the same journal
IEEE Transactions on Automatic Control
Control Engineering

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 68 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf