Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Reducing the Learning Time of Reinforcement Learning for the Supervisory Control of Discrete Event Systems
School of Electro-Mechanical Engineering, Xidian University, Xi’an, China.ORCID-id: 0000-0002-5988-0335
KTH, Skolan för industriell teknik och management (ITM), Maskinkonstruktion, Mekatronik och inbyggda styrsystem.ORCID-id: 0000-0003-4535-3849
KTH, Skolan för industriell teknik och management (ITM), Maskinkonstruktion, Mekatronik och inbyggda styrsystem.ORCID-id: 0000-0001-5703-5923
Department of Industrial Engineering, College of Engineering, King Saud University, Riyadh, Saudi Arabia.ORCID-id: 0000-0003-3559-6249
Vise andre og tillknytning
2023 (engelsk)Inngår i: IEEE Access, E-ISSN 2169-3536, Vol. 11, s. 59840-59853Artikkel i tidsskrift (Fagfellevurdert) Published
Abstract [en]

Reinforcement learning (RL) can obtain the supervisory controller for discrete-event systems modeled by finite automata and temporal logic. The published methods often have two limitations. First, a large number of training data are required to learn the RL controller. Second, the RL algorithms do not consider uncontrollable events, which are essential for supervisory control theory (SCT). To address the limitations, we first apply SCT to find the supervisors for the specifications modeled by automata. These supervisors remove illegal training data violating these specifications and hence reduce the exploration space of the RL algorithm. For the remaining specifications modeled by temporal logic, the RL algorithm is applied to search for the optimal control decision within the confined exploration space. Uncontrollable events are considered by the RL algorithm as uncertainties in the plant model. The proposed method can obtain a nonblocking supervisor for all specifications with less learning time than the published methods.

sted, utgiver, år, opplag, sider
IEEE, 2023. Vol. 11, s. 59840-59853
Emneord [en]
Discrete event system, linear temporal logic, supervisory control theory, reinforcement learning
HSV kategori
Forskningsprogram
Tillämpad matematik och beräkningsmatematik, Optimeringslära och systemteori; Datalogi; Industriella informations- och styrsystem
Identifikatorer
URN: urn:nbn:se:kth:diva-330695DOI: 10.1109/access.2023.3285432ISI: 001018594800001Scopus ID: 2-s2.0-85163172875OAI: oai:DiVA.org:kth-330695DiVA, id: diva2:1778207
Prosjekter
XPRES
Forskningsfinansiär
XPRES - Initiative for excellence in production research
Merknad

QC 20230704

Tilgjengelig fra: 2023-06-30 Laget: 2023-06-30 Sist oppdatert: 2023-07-13bibliografisk kontrollert

Open Access i DiVA

Fulltekst mangler i DiVA

Andre lenker

Forlagets fulltekstScopushttps://ieeexplore.ieee.org/abstract/document/10149832/authors#authors

Person

Tan, KaigeFeng, Lei

Søk i DiVA

Av forfatter/redaktør
Yang, JunjunTan, KaigeFeng, LeiEl-Sherbeeny, Ahmed M.Li, Zhiwu
Av organisasjonen
I samme tidsskrift
IEEE Access

Søk utenfor DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric

doi
urn-nbn
Totalt: 326 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf