kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Adaptive Reinforcement Learning for Real-World Systems with Delays
KTH, School of Electrical Engineering and Computer Science (EECS).
2024 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Reinforcement Learning (RL) allows for learning optimal control strategies for sequential decision-making in processes using the framework of Markov Decision Process (MDP). Regular MDPs assume that the actions selected at a given time step are executed instantaneously and the agent has access to the full state and reward resulting from those actions. However, in realistic scenarios the agent will usually observe the result of the selected actions only after some time, due to, for example, communication lags or response times of physical systems. While the highest focus in existing literature lies on handling known and constant delays, the framework introduced in this thesis allows for an extension to unknown bounded delays. Instead of controlling the process by a single entity, we define a few operational modes and their respective agents. With full knowledge about delay, it is possible to use a delay-informed agent. However, that needs to be preceded by a delay identification period. To satisfy the safety requirements in real-world applications, in this mode the command is handed over to a delay-safe agent, whose purpose is to stabilise the process with robustness to any delay within the considered range. With a component able to manage the transition between those operation modes and corresponding agents, the control framework can autonomously adapt to any value of delay present in the true process. Additionally, the proposed solution uses pre-trained components, consequently minimising the training effort and exploration in the real process, which is crucial for industrial applications.

Abstract [sv]

Förstärkningsinlärning (RL) gör det möjligt att lära sig optimala reglerstrategier för sekventiellt beslutsfattande i processer modellerade som markovbeslutsprocesser (MDP). Klassiska MDPer förutsätter att de handlingar som väljs vid ett givet tidssteg utförs omedelbart och att agenten har tillgång till information om hela tillståndsvektorn och den belöning som följer av dessa handlingar. I realistiska scenarier kommer dock agenten vanligtvis att erhålla information om resultatet av de valda handlingarna först efter en viss tid, till exempel på grund av kommunikationsfördröjningar eller svarstider för fysiska system. I den befintliga litteraturen ligger fokus på att hantera kända och konstanta tidsfördröjningar, men det ramverk som introduceras i den här rapporten möjliggör en utvidgning till okända, begränsade tidsfördröjningar. Istället för att styra processen med en enda enhet definierar vi ett fåtal driftlägen och deras motsvarande agenter. Med full kunskap om tidsfördröjningen är det möjligt att använda en speciell fördröjningsinformerad agent. Detta måste dock föregås av en period där tidsfördröjningens storlek skattas. För att uppfylla säkerhetskraven i verkliga tillämpningar överlämnas kommandot i detta läge till en annan agent, vars syfte är att stabilisera processen med robusthet mot alla fördröjningar inom det aktuella intervallet. Kompletterad med en komponent som kan hantera övergången mellan dessa driftlägen och motsvarande agenter kan reglerramverket autonomt anpassa sig till alla storlekar på tidsfördröjningen som tros kunna uppstå i den verkliga processen. Utöver detta använder den föreslagna metoden förtränade komponenter. Detta minimerar tiden spenderad på träning och anpassning till den verkliga processen vilket är av stor vikt för industriella tillämpningar.

Place, publisher, year, edition, pages
2024. , p. 90
Series
TRITA-EECS-EX ; 2024:830
Keywords [en]
reinforcement learning, delays, actor-critic
Keywords [sv]
förstärkningsinlärning, tidsfördröjningar, aktör-kritiker
National Category
Electrical Engineering, Electronic Engineering, Information Engineering
Identifiers
URN: urn:nbn:se:kth:diva-360480OAI: oai:DiVA.org:kth-360480DiVA, id: diva2:1940424
External cooperation
Västerås, ABB AB
Supervisors
Examiners
Available from: 2025-02-27 Created: 2025-02-26 Last updated: 2025-02-27Bibliographically approved

Open Access in DiVA

fulltext(10390 kB)416 downloads
File information
File name FULLTEXT01.pdfFile size 10390 kBChecksum SHA-512
4d3c32c7179979e58e7327d7f9db72596d476dc15ea13086717005cb5fcffe3e0fa6f972b88ccedc782af9dd4762841fd2cbfaf3df5f833130ddef6827e57f21
Type fulltextMimetype application/pdf

By organisation
School of Electrical Engineering and Computer Science (EECS)
Electrical Engineering, Electronic Engineering, Information Engineering

Search outside of DiVA

GoogleGoogle Scholar
Total: 416 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 862 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf