kth.sePublikationer KTH
Ändra sökning
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Synergies between Policy Learning and Sampling-based Planning
KTH, Skolan för elektroteknik och datavetenskap (EECS), Intelligenta system, Robotik, perception och lärande, RPL.ORCID-id: 0000-0002-1772-7930
2024 (Engelska)Doktorsavhandling, sammanläggning (Övrigt vetenskapligt)Alternativ titel
Synergier mellan policyinlärning och sampling-baserad planering (Svenska)
Abstract [en]

Recent advances in artificial intelligence and machine learning have significantly impacted the field of robotics and led to the interdisciplinary study of robot learning. These developments have the potential to revolutionize the automation of tasks in various industries by reducing the reliance on human workers. However, fully autonomous, learning-based robotic systems are still mainly limited to controlled environments. Ideally, we are looking for methods that enable autonomous acquisition of robotic skills for any temporally extended setting with potentially complex sensor observations. Classical sampling-based planning algorithms used in robot motion planning compute feasible paths between robot states over long time horizons and even in geometrically complex environments. This thesis investigates the possibility of combining learning-based methods with these classical approaches to solve challenging problems in robot manipulation, e.g. the manipulation of deformable objects. The core idea is to leverage the best of both worlds and achieve long-horizon control through planning, while using learning to obtain useful environment models from potentially high-dimensional and complex observation data. The presented frameworks rely on recent machine learning techniques such as contrastive representation learning, generative modeling and reinforcement learning. Finally, we outline the potentials, challenges and limitations of this type of approaches and highlight future directions.

Abstract [sv]

De senaste framstegen inom artificiell intelligens och maskininlärning har haft en betydande inverkan på robotikområdet och lett till det tvärvetenskapliga studerandet av robotinlärning. Dessa utvecklingar har potentialen att revolutionera automatiseringen inom olika industrier genom att minska beroendet av mänskliga arbetare. Dock är helt autonoma, inlärningsbaserade robotsystem fortfarande huvudsakligen begränsade till kontrollerade miljöer. Idealt sett letar vi efter metoder som möjliggör autonom förvärvning av robotfärdigheter för situationer med långa tidshorisonter och potentiellt komplexa sensorobservationer. Klassiska sampling-baserade planeringsalgoritmer som används i robotrörelseplanering beräknar genomförbara vägar mellan robottillstånd över långa tidshorisonter och även i geometriskt komplexa miljöer. I detta arbete undersöker vi möjligheten att kombinera inlärningsbaserade tillvägagångssätt med dessa klassiska tillvägagångssätt för att lösa utmanande problem inom robotmanipulation, t.ex. hantering av formbara objekt. Kärnidén är att utnyttja det bästa av båda världarna och uppnå långsiktig kontroll genom planering, samtidigt som man använder inlärning för att erhålla användbara miljömodeller från potentiellt högdimensionella och komplexa observationsdata. De presenterade ramverken förlitar sig på senaste maskininlärningstekniker såsom kontrastiv representationsinlärning, generativ modellering och förstärkningsinlärning. Slutligen skisserar vi potentialerna, utmaningarna och begränsningarna med denna typ av tillvägagångssätt och belyser framtida riktningar.

Ort, förlag, år, upplaga, sidor
Stockholm, Sweden: KTH Royal Institute of Technology, 2024. , s. ix, 54
Serie
TRITA-EECS-AVL ; 2024:6
Nyckelord [en]
Machine Learning, Robotics, Reinforcement Learning, Motion Planning, Robotic Manipulation
Nationell ämneskategori
Datorgrafik och datorseende
Forskningsämne
Datalogi
Identifikatorer
URN: urn:nbn:se:kth:diva-341911ISBN: 978-91-8040-803-5 (tryckt)OAI: oai:DiVA.org:kth-341911DiVA, id: diva2:1824523
Disputation
2024-01-30, https://kth-se.zoom.us/j/63888939859, F3 (Flodis), Lindstedtsvägen 26 & 28, Stockholm, 15:00 (Engelska)
Opponent
Handledare
Anmärkning

QC 20240108

Tillgänglig från: 2024-01-08 Skapad: 2024-01-05 Senast uppdaterad: 2025-02-07Bibliografiskt granskad
Delarbeten
1. ReForm: A Robot Learning Sandbox for Deformable Linear Object Manipulation
Öppna denna publikation i ny flik eller fönster >>ReForm: A Robot Learning Sandbox for Deformable Linear Object Manipulation
2021 (Engelska)Ingår i: 2021 IEEE International Conference on Robotics and Automation (ICRA), Institute of Electrical and Electronics Engineers (IEEE) , 2021, s. 4717-4723Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Recent advances in machine learning have triggered an enormous interest in using learning-based approaches for robot control and object manipulation. While the majority of existing algorithms are evaluated under the assumption that the involved bodies are rigid, a large number of practical applications contain deformable objects. In this work we focus on Deformable Linear Objects (DLOs) which can be used to model cables, tubes or wires. They are present in many applications such as manufacturing, agriculture and medicine. New methods in robotic manipulation research are often demonstrated in custom environments impeding reproducibility and comparisons of algorithms. We introduce ReForm, a simulation sandbox and a tool for benchmarking manipulation of DLOs. We offer six distinct environments representing important characteristics of deformable objects such as elasticity, plasticity or self-collisions and occlusions. A modular framework is used, enabling design parameters such as the end-effector degrees of freedom, reward function and type of observation. ReForm is a novel robot learning sandbox with which we intend to facilitate testing and reproducibility in manipulation research for DLOs.

Ort, förlag, år, upplaga, sidor
Institute of Electrical and Electronics Engineers (IEEE), 2021
Serie
IEEE International Conference on Robotics and Automation ICRA, ISSN 1050-4729
Nationell ämneskategori
Robotik och automation
Identifikatorer
urn:nbn:se:kth:diva-311663 (URN)10.1109/ICRA48506.2021.9561766 (DOI)000765738803098 ()2-s2.0-85116802097 (Scopus ID)
Konferens
IEEE International Conference on Robotics and Automation (ICRA), MAY 30-JUN 05, 2021, Xian, China
Anmärkning

Part of proceedings: ISBN 978-1-7281-9077-8

QC 20220503

Tillgänglig från: 2022-05-03 Skapad: 2022-05-03 Senast uppdaterad: 2025-02-09Bibliografiskt granskad
2. Planning-Augmented Hierarchical Reinforcement Learning
Öppna denna publikation i ny flik eller fönster >>Planning-Augmented Hierarchical Reinforcement Learning
2021 (Engelska)Ingår i: IEEE Robotics and Automation Letters, E-ISSN 2377-3766, Vol. 6, nr 3, s. 5097-5104Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

Planning algorithms are powerful at solving long-horizon decision-making problems but require that environment dynamics are known. Model-free reinforcement learning has recently been merged with graph-based planning to increase the robustness of trained policies in state-space navigation problems. Recent ideas suggest to use planning in order to provide intermediate waypoints guiding the policy in long-horizon tasks. Yet, it is not always practical to describe a problem in the setting of state-to-state navigation. Often, the goal is defined by one or multiple disjoint sets of valid states or implicitly using an abstract task description. Building upon previous efforts, we introduce a novel algorithm called Planning-Augmented Hierarchical Reinforcement Learning (PAHRL) which translates the concept of hybrid planning/RL to such problems with implicitly defined goal. Using a hierarchical framework, we divide the original task, formulated as a Markov Decision Process (MDP), into a hierarchy of shorter horizon MDPs. Actor-critic agents are trained in parallel for each level of the hierarchy. During testing, a planner then determines useful subgoals on a state graph constructed at the bottom level of the hierarchy. The effectiveness of our approach is demonstrated for a set of continuous control problems in simulation including robot arm reaching tasks and the manipulation of a deformable object.

Ort, förlag, år, upplaga, sidor
Institute of Electrical and Electronics Engineers (IEEE), 2021
Nyckelord
Machine learning for robot control, Motion and path planning, Reinforcement learning
Nationell ämneskategori
Datavetenskap (datalogi)
Identifikatorer
urn:nbn:se:kth:diva-295823 (URN)10.1109/LRA.2021.3071062 (DOI)000642765100014 ()2-s2.0-85103879789 (Scopus ID)
Anmärkning

QC 20210602

Tillgänglig från: 2021-06-02 Skapad: 2021-06-02 Senast uppdaterad: 2024-03-18Bibliografiskt granskad
3. Latent Planning via Expansive Tree Search
Öppna denna publikation i ny flik eller fönster >>Latent Planning via Expansive Tree Search
2022 (Engelska)Ingår i: Advances in Neural Information Processing Systems 35 - 36th Conference on Neural Information Processing Systems, NeurIPS 2022, Neural Information Processing Systems Foundation , 2022Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Planning enables autonomous agents to solve complex decision-making problems by evaluating predictions of the future. However, classical planning algorithms often become infeasible in real-world settings where state spaces are high-dimensional and transition dynamics unknown. The idea behind latent planning is to simplify the decision-making task by mapping it to a lower-dimensional embedding space. Common latent planning strategies are based on trajectory optimization techniques such as shooting or collocation, which are prone to failure in long-horizon and highly non-convex settings. In this work, we study long-horizon goal-reaching scenarios from visual inputs and formulate latent planning as an explorative tree search. Inspired by classical sampling-based motion planning algorithms, we design a method which iteratively grows and optimizes a tree representation of visited areas of the latent space. To encourage fast exploration, the sampling of new states is biased towards sparsely represented regions within the estimated data support. Our method, called Expansive Latent Space Trees (ELAST), relies on self-supervised training via contrastive learning to obtain (a) a latent state representation and (b) a latent transition density model. We embed ELAST into a model-predictive control scheme and demonstrate significant performance improvements compared to existing baselines given challenging visual control tasks in simulation, including the navigation for a deformable object.

Ort, förlag, år, upplaga, sidor
Neural Information Processing Systems Foundation, 2022
Serie
Advances in Neural Information Processing Systems, ISSN 1049-5258 ; 35
Nationell ämneskategori
Robotik och automation Datorgrafik och datorseende
Identifikatorer
urn:nbn:se:kth:diva-331664 (URN)2-s2.0-85163176952 (Scopus ID)
Konferens
36th Conference on Neural Information Processing Systems, NeurIPS 2022, New Orleans, United States of America, Nov 28 2022 - Dec 9 2022
Anmärkning

Part of ISBN 9781713871088

QC 20230712

Tillgänglig från: 2023-07-13 Skapad: 2023-07-13 Senast uppdaterad: 2025-02-05Bibliografiskt granskad
4. Expansive Latent Planning for Sparse Reward Offline Reinforcement Learning
Öppna denna publikation i ny flik eller fönster >>Expansive Latent Planning for Sparse Reward Offline Reinforcement Learning
2023 (Engelska)Ingår i: Proceedings of The 7th Conference on Robot Learning, Proceedings of Machine Learning Research , 2023Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Sampling-based motion planning algorithms excel at searching global solution paths in geometrically complex settings. However, classical approaches, such as RRT, are difficult to scale beyond low-dimensional search spaces and rely on privileged knowledge e.g. about collision detection and underlying state distances. In this work, we take a step towards the integration of sampling-based planning into the reinforcement learning framework to solve sparse-reward control tasks from high-dimensional inputs. Our method, called VELAP, determines sequences of waypoints through sampling-based exploration in a learned state embedding. Unlike other sampling-based techniques, we iteratively expand a tree-based memory of visited latent areas, which is leveraged to explore a larger portion of the latent space for a given number of search iterations. We demonstrate state-of-the-art results in learning control from offline data in the context of vision-based manipulation under sparse reward feedback. Our method extends the set of available planning tools in model-based reinforcement learning by adding a latent planner that searches globally for feasible paths instead of being bound to a fixed prediction horizon. 

Ort, förlag, år, upplaga, sidor
Proceedings of Machine Learning Research, 2023
Nationell ämneskategori
Datorgrafik och datorseende Robotik och automation
Identifikatorer
urn:nbn:se:kth:diva-341581 (URN)001221201500001 ()2-s2.0-85184350420 (Scopus ID)
Konferens
The 7th Conference on Robot Learning, Atlanta, GA, Nov 6-9, 2023
Anmärkning

QC 20231227

Tillgänglig från: 2023-12-22 Skapad: 2023-12-22 Senast uppdaterad: 2025-02-05Bibliografiskt granskad

Open Access i DiVA

Kappa(4021 kB)1042 nedladdningar
Filinformation
Filnamn FULLTEXT01.pdfFilstorlek 4021 kBChecksumma SHA-512
1a6573a890731d099cc1c6b63dc688404aa2cf2fbbe8ead8519fe496d3c662cc36341d0af8b4b8344561f4c3b34226a528d6e9315933cd7a561d10102a3383f4
Typ fulltextMimetyp application/pdf

Person

Gieselmann, Robert

Sök vidare i DiVA

Av författaren/redaktören
Gieselmann, Robert
Av organisationen
Robotik, perception och lärande, RPL
Datorgrafik och datorseende

Sök vidare utanför DiVA

GoogleGoogle Scholar
Totalt: 1043 nedladdningar
Antalet nedladdningar är summan av nedladdningar för alla fulltexter. Det kan inkludera t.ex tidigare versioner som nu inte längre är tillgängliga.

isbn
urn-nbn

Altmetricpoäng

isbn
urn-nbn
Totalt: 3109 träffar
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf