kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Publications (4 of 4) Show all publications
Hjelm, M. (2019). Holistic Grasping: Affordances, Grasp Semantics, Task Constraints. (Doctoral dissertation). Stockholm: KTH Royal Institute of Technology
Open this publication in new window or tab >>Holistic Grasping: Affordances, Grasp Semantics, Task Constraints
2019 (English)Doctoral thesis, monograph (Other academic)
Abstract [en]

Most of us perform grasping actions over a thousand times per day without giving it much consideration, be it from driving to drinking coffee. Learning robots the same ease when it comes to grasping has been a goal for the robotics research community for decades.

The reason for the slow progress lays mainly in the inferiority of the robot sensorimotor system. Robotic grippers are often non-compliant, lack the degrees of freedom of human hands, and haptic sensors are rudimentary involving significantly less resolution and sensitivity than in humans.

Research has therefore focused on engineering solutions that center on the stability of the grasp. This involves specifying complex functions and search strategies detailing the interaction between the digits of the robot and the surface of the object. Given the amount of variation in materials, shapes, and ability to deform it seems infeasible to analytically formulate such a gripper-to-shape mapping. Many researchers have instead looked to data-driven methods for learning the gripper-to-shape mapping as does this thesis.

Humans obviously have a similar mapping capability. However, how we grasp an object is determined foremost by what we are going to do with the object. We have priors on task, material, and the dynamics of objects that help guide the grasping process. We also have a deeper understanding of how shape and material relate to our own embodiment.

We tie all these aspects together: our understanding of what an object can be used for, how that affects our interaction with it, and how our hand can form to achieve the goal of the manipulation. For us humans grasping is not just a gripper-to-shape mapping it is a holistic process where all parts of the chain matters to the outcome. The focus of this thesis is thus on how to incorporate such a holistic process into robotic grasp planning.  

We will address the holistic grasping process through three jointly connected modules. The first is affordance detection and learning to infer the common parts for objects that afford an action, a form of conceptualization of the affordance categories. The second is learning grasp semantics, how shape relates to the gripper configuration. And finally the third is to learn how task constrains the grasping process.

We will explore these three parts through the concept of similarity. This translates directly into the idea that we should learn a representation that puts similar types of the entities that we are describing, that is, objects, grasps, and tasks, close to each other in space. We will show that the idea of similarity based representations will help the robot reason about which parts of an object is important for affordance inference, which grasps and tasks are similar, and how the categories relate to each other. Finally, the similarity-based approach will help us tie all parts together in the conceptual demonstration of how a holistic grasping process might be realized.

Abstract [sv]

De flesta av oss greppar objekt över tusen gånger per dag utan att ge det mycket eftertanke, vare sig det är att köra bil eller att dricka kaffe. Att lära robotar liknande förmågor gällande manipulering har varit ett mål för robotforskningen i årtionden.

Anledningen till de långsamma framstegen ligger huvudsakligen i robotarnas underutvecklade sensorimotoriska system. Robothänder är ofta inflexibla, saknar möjligheter till komplexa konfigurationer jämfört med mänskliga händer. De haptiska sensorerna är rudimentära, vilket innebär betydligt lägre upplösning och känslighet vid beröring än hos människor.

Den nuvarande forskningen har därför koncentrerat sig på tekniska lösningar som fokuserar på stabiliteten i det slutgiltiga greppet. Detta innebär att man formulerar komplexa funktioner och sökstrategier som beskriver interaktionen mellan robotens fingar och objektets yta. Med tanke på mängden variation i material, former och förmåga att deformera verkar det otänkbart att kunna analytiskt formulera en sådan generell hand-till-form-funktion. Många forskare har istället börjat fokusera på metoder baserade på lärande från data, likså den här avhandlingen.

Människor har uppenbarligen en förmåga att synka hand till form. Hur vi greppar ett objekt bestäms emellertid främst av vad vi ska göra med objektet. Vi har en intern a priori uppfattning av hur handlingen, material och objektdynamiken styr grepp-processen. Vi har också en djupare förståelse för hur form och material relaterar till vår egen hand.

Vi knyter samman alla dessa aspekter: vår förståelse för vad ett föremål kan användas för, hur den användningen påverkar vår interaktion med det och hur vår hand kan formas och placeras för att uppnå målet för manipulationen. För oss är grepp-processen inte bara en hand-till-form funktion utan en holistisk process där alla delar av kedjan är lika viktiga för resultatet. Innehållet i denna avhandling handlar således om hur man införlivar en sådan process i en robots planering av maipulationsmomentet.

Vi kommer ta oss an den holistiska processen genom tre sammankopplade moduler. Den första är att låta roboten detektera interaktionsmöjligheter och förstå vilka delar av ett objekt som är viktiga för att möjliggöra interaktionen, en form a konceptualisering av interaktionsmöjligheten. Den andra modulen handlar om utlärning av grepp semantik, hur form relaterar till den egna handens  förmåga. Slutligen är sista modulen fokuserad på hur man lär roboten hur målet med interaktionen påverkar möjliga grepp på objektet.Vi kommer att utforska dessa tre delar genom begreppet affinitet. Detta begrepp translateras direkt till idén att vi lär oss en representation som sätter liknande typer av entiteter, det vill säga objekt, grepp, och mål, nära varandra i representationsrymden.

Vi kommer att visa att idén om affinitetsbaserade representationer kommer att hjälpa roboten a resonera kring vilka delar av ett objekt som är viktiga för inferens, vilka grepp och mål som liknar varandra och hur de olika kategorierna relaterar till varandra. Slutligen kommer ett affinitetsbaserat tillvägagångssätt att hjälpa oss att knyta samman alla delar i en demonstrationen av en holistisk grepp-process.

Place, publisher, year, edition, pages
Stockholm: KTH Royal Institute of Technology, 2019. p. 178
Series
TRITA-EECS-AVL ; 2019:48
Keywords
robotics, robotic grasping, grasping, cognition, embodied cognition, computer vision, machine learning, artificial intelligence, AI, Gaussian process, Gaussian process latent variable model, GPLVM, 3D vision, point cloud features, robotik, manipulation, datorseende, maskininlärning, artificiell intelligens, kognition
National Category
Computer graphics and computer vision
Research subject
Computer Science
Identifiers
urn:nbn:se:kth:diva-251388 (URN)978-91-7873-203-6 (ISBN)
Public defence
2019-06-04, D2, Lindstedtsvägen 5, 114 28 Stockholm, Stockholm, 10:00 (English)
Opponent
Supervisors
Note

QC20190514

Available from: 2019-05-14 Created: 2019-05-13 Last updated: 2025-02-07Bibliographically approved
Hjelm, M., Ek, C. H., Detry, R. & Kragic, D. (2015). Learning Human Priors for Task-Constrained Grasping. In: COMPUTER VISION SYSTEMS (ICVS 2015): . Paper presented at 10th International Conference on Computer Vision Systems (ICVS), JUL 06-09, 2015, Copenhagen, DENMARK (pp. 207-217). Springer Berlin/Heidelberg
Open this publication in new window or tab >>Learning Human Priors for Task-Constrained Grasping
2015 (English)In: COMPUTER VISION SYSTEMS (ICVS 2015), Springer Berlin/Heidelberg, 2015, p. 207-217Conference paper, Published paper (Refereed)
Abstract [en]

An autonomous agent using manmade objects must understand how task conditions the grasp placement. In this paper we formulate task based robotic grasping as a feature learning problem. Using a human demonstrator to provide examples of grasps associated with a specific task, we learn a representation, such that similarity in task is reflected by similarity in feature. The learned representation discards parts of the sensory input that is redundant for the task, allowing the agent to ground and reason about the relevant features for the task. Synthesized grasps for an observed task on previously unseen objects can then be filtered and ordered by matching to learned instances without the need of an analytically formulated metric. We show on a real robot how our approach is able to utilize the learned representation to synthesize and perform valid task specific grasps on novel objects.

Place, publisher, year, edition, pages
Springer Berlin/Heidelberg, 2015
Series
Lecture Notes in Computer Science, ISSN 0302-9743 ; 9163
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-177975 (URN)10.1007/978-3-319-20904-3_20 (DOI)000364183300020 ()2-s2.0-84949035044 (Scopus ID)978-3-319-20904-3 (ISBN)978-3-319-20903-6 (ISBN)
Conference
10th International Conference on Computer Vision Systems (ICVS), JUL 06-09, 2015, Copenhagen, DENMARK
Note

QC 20151202

Available from: 2015-12-02 Created: 2015-11-30 Last updated: 2025-02-07Bibliographically approved
Hjelm, M., Detry, R., Ek, C. H. & Kragic, D. (2014). Representations for cross-task, cross-object grasp Transfer. In: Proceedings - IEEE International Conference on Robotics and Automation: . Paper presented at 2014 IEEE International Conference on Robotics and Automation, ICRA 2014, 31 May 2014 through 7 June 2014 (pp. 5699-5704). IEEE conference proceedings
Open this publication in new window or tab >>Representations for cross-task, cross-object grasp Transfer
2014 (English)In: Proceedings - IEEE International Conference on Robotics and Automation, IEEE conference proceedings, 2014, p. 5699-5704Conference paper, Published paper (Refereed)
Abstract [en]

We address The problem of Transferring grasp knowledge across objects and Tasks. This means dealing with Two important issues: 1) The induction of possible Transfers, i.e., whether a given object affords a given Task, and 2) The planning of a grasp That will allow The robot To fulfill The Task. The induction of object affordances is approached by abstracting The sensory input of an object as a set of attributes That The agent can reason about Through similarity and proximity. For grasp execution, we combine a part-based grasp planner with a model of Task constraints. The Task constraint model indicates areas of The object That The robot can grasp To execute The Task. Within These areas, The part-based planner finds a hand placement That is compatible with The object shape. The key contribution is The ability To Transfer Task parameters across objects while The part-based grasp planner allows for Transferring grasp information across Tasks. As a result, The robot is able To synthesize plans for previously unobserved Task/object combinations. We illustrate our approach with experiments conducted on a real robot.

Place, publisher, year, edition, pages
IEEE conference proceedings, 2014
National Category
Electrical Engineering, Electronic Engineering, Information Engineering
Identifiers
urn:nbn:se:kth:diva-176152 (URN)10.1109/ICRA.2014.6907697 (DOI)000377221105111 ()2-s2.0-84929192413 (Scopus ID)
Conference
2014 IEEE International Conference on Robotics and Automation, ICRA 2014, 31 May 2014 through 7 June 2014
Note

QC 20151130

Available from: 2015-11-30 Created: 2015-11-02 Last updated: 2024-03-15Bibliographically approved
Hjelm, M., Ek, C. H., Detry, R., Kjellström, H. & Kragic, D. (2013). Sparse Summarization of Robotic Grasping Data. In: 2013 IEEE International Conference on Robotics and Automation (ICRA): . Paper presented at 2013 IEEE International Conference on Robotics and Automation, ICRA 2013; Karlsruhe; Germany; 6 May 2013 through 10 May 2013 (pp. 1082-1087). New York: IEEE
Open this publication in new window or tab >>Sparse Summarization of Robotic Grasping Data
Show others...
2013 (English)In: 2013 IEEE International Conference on Robotics and Automation (ICRA), New York: IEEE , 2013, p. 1082-1087Conference paper, Published paper (Refereed)
Abstract [en]

We propose a new approach for learning a summarized representation of high dimensional continuous data. Our technique consists of a Bayesian non-parametric model capable of encoding high-dimensional data from complex distributions using a sparse summarization. Specifically, the method marries techniques from probabilistic dimensionality reduction and clustering. We apply the model to learn efficient representations of grasping data for two robotic scenarios.

Place, publisher, year, edition, pages
New York: IEEE, 2013
Series
IEEE International Conference on Robotics and Automation, ISSN 1050-4729
Keywords
Principal Component Analysis, Models
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-136187 (URN)10.1109/ICRA.2013.6630707 (DOI)000337617301013 ()2-s2.0-84887316002 (Scopus ID)978-1-4673-5643-5 (ISBN)978-1-4673-5641-1 (ISBN)
Conference
2013 IEEE International Conference on Robotics and Automation, ICRA 2013; Karlsruhe; Germany; 6 May 2013 through 10 May 2013
Note

QC 20140129

Available from: 2013-12-04 Created: 2013-12-04 Last updated: 2025-02-07Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-1031-9600

Search in DiVA

Show all publications