Open this publication in new window or tab >>Show others...
2027 (English)In: Robotics and Computer-Integrated Manufacturing, ISSN 0736-5845, E-ISSN 1879-2537, Vol. 103, article id 103382Article in journal (Refereed) Published
Abstract [en]
The advancement of Industry 5.0 has positioned human-robot collaboration (HRC) as a critical component of smart manufacturing, requiring robots to work seamlessly alongside humans in dynamic, unstructured environments. However, several challenges remain, including the heavy reliance on object-specific data for training dynamic perception models, the reliance on manually prepared CAD models for 6D pose estimation, and the lack of interaction semantics in grasp planning. To this end, a one-shot human-centric robotic grasping framework that does not require object-specific training or manually prepared CAD models, namely OH-Grasp, is proposed for HRC scenarios. This framework uses a single masked reference RGB-D observation to support instance segmentation, 6D pose estimation, and human-centric grasp planning for query RGB-D observations. First, the one-shot learning-based instance segmentation (OSeg) module is proposed without object-specific training, which utilizes the semantic consistency and geometric constraints of pre-trained visual encoders to guide a promptable segmentation foundation model, achieving robust one-shot instance segmentation. Second, the one-shot learning-based 6D pose estimation (OPose) module is proposed without manually prepared CAD models, combining single-view 3D reconstruction with differentiable rendering optimization to address the challenges of scale and 6D pose estimation for new objects. Finally, the human-centric robotic grasping (HGrasp) module is proposed without object-specific training, which leverages Vision-Language Models (VLMs) to parse object functional regions and integrates them with 3D geometric constraints and a geometry-center prior to generate grasping strategies that align with human interaction habits. Real-robot experiments on 15 tools, 5 parts, and three multi-object scenes show that OH-Grasp achieves grasping quality values of 95.89%, 89.92%, and 95.16%, with an overall end-to-end system success rate of 91.11% over 180 trials. Under conservative safety settings, the scene-level execution times are 25.91 s, 29.15 s, and 32.24 s. A human-centric handover evaluation with 12 participants and 28,800 pairwise Likert ratings further shows that OH-Grasp is preferred in 20 of 24 Directness and Security comparisons. These results demonstrate the effectiveness of OH-Grasp for task-level one-shot human-centric grasping in dynamic and partially unstructured HRC scenarios.
Place, publisher, year, edition, pages
Elsevier BV, 2027
Keywords
Human-robot collaboration, One-shot learning, Instance segmentation, 6D pose estimation, Human-centric grasping
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-387729 (URN)10.1016/j.rcim.2026.103382 (DOI)001823545800001 ()2-s2.0-105044109402 (Scopus ID)
Note
QC 20260901
2026-09-012026-09-012026-09-01Bibliographically approved