kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Publications (10 of 55) Show all publications
Zhang, Y., Li, B., Ding, K., Zhang, F., Liu, S., Sun, Y. & Hui, J. (2027). OH-Grasp: A one-shot human-centric robotic grasping framework for human-robot collaboration in Industry 5.0. Robotics and Computer-Integrated Manufacturing, 103, Article ID 103382.
Open this publication in new window or tab >>OH-Grasp: A one-shot human-centric robotic grasping framework for human-robot collaboration in Industry 5.0
Show others...
2027 (English)In: Robotics and Computer-Integrated Manufacturing, ISSN 0736-5845, E-ISSN 1879-2537, Vol. 103, article id 103382Article in journal (Refereed) Published
Abstract [en]

The advancement of Industry 5.0 has positioned human-robot collaboration (HRC) as a critical component of smart manufacturing, requiring robots to work seamlessly alongside humans in dynamic, unstructured environments. However, several challenges remain, including the heavy reliance on object-specific data for training dynamic perception models, the reliance on manually prepared CAD models for 6D pose estimation, and the lack of interaction semantics in grasp planning. To this end, a one-shot human-centric robotic grasping framework that does not require object-specific training or manually prepared CAD models, namely OH-Grasp, is proposed for HRC scenarios. This framework uses a single masked reference RGB-D observation to support instance segmentation, 6D pose estimation, and human-centric grasp planning for query RGB-D observations. First, the one-shot learning-based instance segmentation (OSeg) module is proposed without object-specific training, which utilizes the semantic consistency and geometric constraints of pre-trained visual encoders to guide a promptable segmentation foundation model, achieving robust one-shot instance segmentation. Second, the one-shot learning-based 6D pose estimation (OPose) module is proposed without manually prepared CAD models, combining single-view 3D reconstruction with differentiable rendering optimization to address the challenges of scale and 6D pose estimation for new objects. Finally, the human-centric robotic grasping (HGrasp) module is proposed without object-specific training, which leverages Vision-Language Models (VLMs) to parse object functional regions and integrates them with 3D geometric constraints and a geometry-center prior to generate grasping strategies that align with human interaction habits. Real-robot experiments on 15 tools, 5 parts, and three multi-object scenes show that OH-Grasp achieves grasping quality values of 95.89%, 89.92%, and 95.16%, with an overall end-to-end system success rate of 91.11% over 180 trials. Under conservative safety settings, the scene-level execution times are 25.91 s, 29.15 s, and 32.24 s. A human-centric handover evaluation with 12 participants and 28,800 pairwise Likert ratings further shows that OH-Grasp is preferred in 20 of 24 Directness and Security comparisons. These results demonstrate the effectiveness of OH-Grasp for task-level one-shot human-centric grasping in dynamic and partially unstructured HRC scenarios.

Place, publisher, year, edition, pages
Elsevier BV, 2027
Keywords
Human-robot collaboration, One-shot learning, Instance segmentation, 6D pose estimation, Human-centric grasping
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-387729 (URN)10.1016/j.rcim.2026.103382 (DOI)001823545800001 ()2-s2.0-105044109402 (Scopus ID)
Note

QC 20260901

Available from: 2026-09-01 Created: 2026-09-01 Last updated: 2026-09-01Bibliographically approved
Wang, Q., Liu, Y., Zhu, Z., Liu, S., Wang, Z., Zi, B., . . . Zhang, L. (2026). An MVUN-DP skill library-driven Embodied AI robotic assembly method with LLMs. Robotics and Computer-Integrated Manufacturing, 100, Article ID 103263.
Open this publication in new window or tab >>An MVUN-DP skill library-driven Embodied AI robotic assembly method with LLMs
Show others...
2026 (English)In: Robotics and Computer-Integrated Manufacturing, ISSN 0736-5845, E-ISSN 1879-2537, Vol. 100, article id 103263Article in journal (Refereed) Published
Abstract [en]

Nowadays, multi-variety, small-batch production poses significant challenges to robotic assembly and drives the need for advanced Embodied AI (EAI). However, there are insufficient effective methods for constructing skill libraries that are able to execute assembly of shape-variant components precisely and robustly, and for training assembly-oriented skills as well in terms of EAI. To address these limitations, this paper introduces a MobileViT + U-Net Diffusion Policy (MVUN-DP) skill library-driven assembly framework with a method of constructing diverse assembly skills characterized by high precision and robustness against environmental disturbances. First, a systematic method for constructing a robotic assembly skill library based on the MVUN-DP algorithm is proposed. Designed for objects with diverse geometries, our approach trains assembly skills through domain randomization and dataset augmentation, which significantly enhances robustness against pose uncertainties and environmental variations. Second, the MVUN-DP algorithm is introduced to enhance skill learning capabilities. Building on the expressive action-generation capability of Diffusion Policy (DP), MVUN-DP incorporates a Mobile Vision Transformer (MobileViT) for fine-grained visual feature encoding that enables precise and adaptive motion synthesis. Third, voice-to-motion assembly by seamlessly integrating the MVUN-DP skill library with LLM-based task planning and execution is implemented. The proposed framework is extensively evaluated on a UR5 robotic platform equipped with dual RGB cameras and force/torque sensors. For planetary gear assembly and other shape-variant tasks, MVUN-DP achieves an average success rate of 94.7 %, which outperforms imitation learning baselines in both success rate and composite episode rewards while maintaining low contact forces and torques. Furthermore, data augmentation improves the policy's robustness to environmental disturbances by 7.12 %.

Place, publisher, year, edition, pages
Elsevier BV, 2026
Keywords
Diffusion policy, Imitation learning, LLM, Robot learning, Robotic assembly
National Category
Robotics and automation Production Engineering, Human Work Science and Ergonomics
Identifiers
urn:nbn:se:kth:diva-377328 (URN)10.1016/j.rcim.2026.103263 (DOI)001692130500001 ()2-s2.0-105029748366 (Scopus ID)
Note

QC 20260227

Available from: 2026-02-27 Created: 2026-02-27 Last updated: 2026-02-27Bibliographically approved
Zhao, S., Deng, Z., Liu, S., Xu, H., Zhang, J. & Zhong, R. Y. (2026). CA-DLM: Causal-aware diffusion language model for interpretable industrial defective image generation. Journal of manufacturing systems, 88, 772-785
Open this publication in new window or tab >>CA-DLM: Causal-aware diffusion language model for interpretable industrial defective image generation
Show others...
2026 (English)In: Journal of manufacturing systems, ISSN 0278-6125, E-ISSN 1878-6642, Vol. 88, p. 772-785Article in journal (Refereed) Published
Abstract [en]

This research proposes a Causal-Aware Diffusion Language Model (CA-DLM) for interpretable generation of defective images to address the scarcity of defective samples in industrial settings. Firstly, a causal inference mechanism is integrated into DLM to establish a causal mapping among “natural language instructions, model causal variables, and defective image generation”. The counterfactual causal intervention strategy enables interpretable, fine-grained control over defect, texture, and background features during image generation. Secondly, a learnable Structural Causal Model (SCM) is constructed to disentangle semantic features into causal and non-causal variables, effectively separating useful semantic features from confounders during image generation. A causal loss function is also designed to enforce sparsity and acyclicity in the SCM, guiding the model toward a causally consistent structure. Finally, the causal intervention approach is designed to perform do-intervention on SCM to achieve interpretable defective image generation. The do-intervention activates relevant causal variables and suppresses non-causal variables according to the learned SCM, thereby enabling precise control over key defect attributes while preserving the consistency of texture and background. Comparative experiments across three public datasets show that CA-DLM enables interpretable, controllable generation of high-fidelity defective images, outperforming state-of-the-art methods. Detection models trained on CA-DLM generated images achieve significant improvements in downstream defect detection tasks, with accuracy exceeding 95%. Ablation studies further validate the effectiveness of the proposed approach, showing that high-fidelity defective images can be generated using only 30% of causal variables.

Place, publisher, year, edition, pages
Elsevier BV, 2026
Keywords
Causal Inference, Defective Image Generation, Diffusion Language Model
National Category
Computer Sciences Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-386371 (URN)10.1016/j.jmsy.2026.07.007 (DOI)001830037500001 ()2-s2.0-105044815279 (Scopus ID)
Note

QC 20260803

Available from: 2026-08-03 Created: 2026-08-03 Last updated: 2026-08-03Bibliographically approved
Liu, S., Regal, E., Zhang, H., Wang, C., Wang, L. & Gao, R. X. (2026). Foundation model-based end-to-end policy for robotic assembly. CIRP annals, 75(1), 43-47
Open this publication in new window or tab >>Foundation model-based end-to-end policy for robotic assembly
Show others...
2026 (English)In: CIRP annals, ISSN 0007-8506, E-ISSN 1726-0604, Vol. 75, no 1, p. 43-47Article in journal (Refereed) Published
Abstract [en]

Despite rapid advances in foundation models, achieving robust and reliable performance in complex assembly tasks remains a substantial challenge. This paper introduces a foundation model-based approach for learning end-to-end assembly policies involving both robots and humans. A unified framework built on the LeRobot platform is developed for robot teleoperation, data collection and curation, model fine-tuning, and local deployment on robotic systems. Using assembly-oriented language instructions, state-of-the-art foundation models are fine-tuned under different parameter configurations and evaluated comprehensively. Successful policy deployment with asynchronous inference and long-horizon action chunking on a partial car engine assembly demonstrates the effectiveness of the developed methods.

Place, publisher, year, edition, pages
Elsevier BV, 2026
Keywords
Robot, Assembly, Foundation model
National Category
Robotics and automation
Identifiers
urn:nbn:se:kth:diva-388417 (URN)10.1016/j.cirp.2026.04.104 (DOI)001831726800001 ()2-s2.0-105037120129 (Scopus ID)
Note

QC 20260916

Available from: 2026-09-16 Created: 2026-09-16 Last updated: 2026-09-16Bibliographically approved
Zhao, S., Zhang, G., Liu, S., Zhang, J., Dilum Bandara, H. M., Zhong, R. Y. & Wang, L. (2026). Interpretable Verification Mechanism for Trustworthy Industrial Large Model in Intelligent Manufacturing. Engineering, 63, 84-96
Open this publication in new window or tab >>Interpretable Verification Mechanism for Trustworthy Industrial Large Model in Intelligent Manufacturing
Show others...
2026 (English)In: Engineering, ISSN 2095-8099, Vol. 63, p. 84-96Article in journal (Refereed) Published
Abstract [en]

The hallucination and black-box nature of Large Models limit their industrial applications. To address these challenges, a verification mechanism built on confidence intervals of Transformer-based output layers is proposed for trustworthy Industrial Large Models (ILMs). Adopting a Vision Transformer (ViT), customized verification operations are incorporated to monitor the forward propagation process, and samples with probability distributions outside confidence intervals exit the network early and are handed over to technicians. Thus, the ViT is more interpretable because only samples within confidence intervals can propagate forward and be output from the ViT. Subsequently, an over-approximation approach is employed to obtain confidence intervals by linearizing the decision boundary of the ViT. The conservative decision boundary serves as the lower bound of confidence intervals, which can provide provable robustness for confidence intervals because the minimum probability of the ground truth is always higher than that of other samples. Finally, a certified training strategy is employed to enhance the robustness of the ViT. Data disturbances with Gaussian noise are generated using a randomized smoothing strategy to augment the data distribution. A smoothed loss function is used to strengthen the robustness of the ViT against data disturbances, thereby enabling greater confidence intervals. The proposed verification mechanism was validated on two public defect datasets. It achieved 99.98% precision for normal samples and approximately 95% precision for defective samples on a fabric defect dataset. It also achieved 99.21% precision and 99.15% F1 score on a wafer defect dataset. Comparative experiments with other Transformer-based models also demonstrated the generalization ability of the proposed verification mechanism.

Place, publisher, year, edition, pages
Elsevier BV, 2026
Keywords
Hallucination, Industrial defect detection, Industrial Large Models, Trustworthy
National Category
Artificial Intelligence Robotics and automation
Identifiers
urn:nbn:se:kth:diva-382814 (URN)10.1016/j.eng.2025.08.023 (DOI)2-s2.0-105038618387 (Scopus ID)
Note

QC 20260602

Available from: 2026-06-02 Created: 2026-06-02 Last updated: 2026-08-28Bibliographically approved
Zhu, Z., Liu, Y., Wang, Q., Wang, Z., Wang, L., Liu, S., . . . Zhang, L. (2026). Toward generalizable robotic assembly: A prior-guided deep reinforcement learning approach with multi-sensor information. Robotics and Computer-Integrated Manufacturing, 100, Article ID 103242.
Open this publication in new window or tab >>Toward generalizable robotic assembly: A prior-guided deep reinforcement learning approach with multi-sensor information
Show others...
2026 (English)In: Robotics and Computer-Integrated Manufacturing, ISSN 0736-5845, E-ISSN 1879-2537, Vol. 100, article id 103242Article in journal (Refereed) Published
Abstract [en]

The rise of personalized manufacturing presents significant challenges for robotic assembly. While learning-based methods offer promising solutions, they often suffer from low training efficiency and poor generalization. To address these limitations, this paper proposes an efficient prior-guided (PG) deep reinforcement learning (DRL) approach for generalizable robotic assembly using multi-sensor information. First, a phased multi-sensor information fusion method is introduced. Then, a visual feature extraction method that combines MobileNetV3-Lite with conventional digital image processing and a rule-based force feature extraction method are designed to extract lower-dimensional features as prior-guided knowledge. Based on the methods above, a Soft Actor-Critic (SAC) algorithm that integrates Gated Recurrent Unit (GRU) network architecture with PG is proposed, thereby enabling efficient assembly skill learning. Simulations and physical experiments with respect to three typical assembly skills, i.e., search, alignment, and insertion, are conducted. Results indicate that, compared with the baseline SAC algorithm, our feature extraction method reduces visual feature dimensions by 93.75% and provides accurate prior-guided knowledge for DRL. The proposed assembly skill learning algorithm achieves a 30.16% reduction in average training time and a 16.82% decrease in average completion step. Furthermore, all learned skills can be rapidly transferred across different objects, and all assembly tasks are completed efficiently and compliantly with an average success rate of 96.86%.

Place, publisher, year, edition, pages
Elsevier BV, 2026
Keywords
Deep reinforcement learning, Multi-sensor information fusion, Robot skill learning, Robotic assembly
National Category
Robotics and automation Computer Sciences Computer graphics and computer vision Control Engineering
Identifiers
urn:nbn:se:kth:diva-375996 (URN)10.1016/j.rcim.2026.103242 (DOI)001677206500001 ()2-s2.0-105028023237 (Scopus ID)
Note

QC 20260130

Available from: 2026-01-30 Created: 2026-01-30 Last updated: 2026-05-29Bibliographically approved
Yi, S., Liu, S., Lin, X., Yan, S., Wang, X. V. & Wang, L. (2025). A data-efficient and general-purpose hand–eye calibration method for robotic systems using next best view. Advanced Engineering Informatics, 66, Article ID 103432.
Open this publication in new window or tab >>A data-efficient and general-purpose hand–eye calibration method for robotic systems using next best view
Show others...
2025 (English)In: Advanced Engineering Informatics, ISSN 1474-0346, E-ISSN 1873-5320, Vol. 66, article id 103432Article in journal (Refereed) Published
Abstract [en]

Calibration between robots and cameras is critical in automated robot vision systems. However, conventional manually conducted image-based calibration techniques are often limited by their accuracy sensitivity and poor adaptability to dynamic or unstructured environments. These approaches present challenges for ease of calibration and automatic deployment while being susceptible to rigid assumptions that degrade their performance. To close these limitations, this study proposes a data-efficient vision-driven approach for fast, accurate, and robust hand–eye camera calibration, and it aims to maximise the efficiency of robots in obtaining hand–eye calibration images without compromising accuracy. By analysing the previously captured images, the minimisation of the residual Jacobian matrix is utilised to predict the next optimal pose for robot calibration. A method to adjust the camera poses in dynamic environments is proposed to achieve efficient and robust hand–eye calibration. It requires fewer images, reduces dependence on manual expertise, and ensures repeatability. The proposed method is tested using experiments with actual industrial robots. The results demonstrate that our NBV strategy reduces rotational error by 8.8%, translational error by 26.4%, and the number of sampling frames by 25% compared to artificial sampling. The experimental results show that the average prediction time per frame is 3.26 seconds.

Place, publisher, year, edition, pages
Elsevier BV, 2025
Keywords
Hand–eye calibration, Non-linear optimisation, Robot control, Robot vision system
National Category
Robotics and automation Computer graphics and computer vision Control Engineering
Identifiers
urn:nbn:se:kth:diva-364151 (URN)10.1016/j.aei.2025.103432 (DOI)001504534600004 ()2-s2.0-105005832045 (Scopus ID)
Note

QC 20250605

Available from: 2025-06-04 Created: 2025-06-04 Last updated: 2025-08-15Bibliographically approved
Liu, S., Guo, D., Liu, Z., Wang, T., Qin, Q., Wang, X. V. & Wang, L. (2025). A Digital Twin-Enabled Approach to Reliable Human–robot Collaborative Assembly. In: Human Centric Smart Manufacturing Towards Industry 5 0: (pp. 281-304). Springer Nature
Open this publication in new window or tab >>A Digital Twin-Enabled Approach to Reliable Human–robot Collaborative Assembly
Show others...
2025 (English)In: Human Centric Smart Manufacturing Towards Industry 5 0, Springer Nature , 2025, p. 281-304Chapter in book (Other academic)
Abstract [en]

The conventional automation approach has shown bottlenecks in the era of component assembly. What could be automated has been automated in some high tech industrial production, leaving manual work performed by humans. To achieve ergonomic working environments and better productivity, human–robot collabora tion has been adopted for this purpose through combining the strength, accuracy and repeatability of robots with adaptability, high-level cognition, and flexibility. A reliable human–robot collaborative setting should be supported by dynamically updated and precise models. For this purpose, the digital twin can realise the digital representation of physical collaborative settings through simulation modelling and data synchronisation but is limited by communication delay and constraints. This chapter will develop a digital twin-enabled approach to human–robot collaborative assembly. Within this approach, a sensor-driven 3D modelling of the physical devices of interest is developed to realise the physical-to-digital transformation of human–robot workcell, and a Wise-ShopFloor-based platform enabled by sensor data is used to develop a digital twin model of the physical human–robot workcell. Then, function blocks with embedded algorithms are used for assembly planning, decision making and robot control, and a time-ahead execution and planning approach is developed for reliable human–robot collaborative assembly. Finally, the performance of the developed system is demonstrated by a case study of a partial car engine assembly.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
Assembly, Digital twin, Robot
National Category
Production Engineering, Human Work Science and Ergonomics
Identifiers
urn:nbn:se:kth:diva-368723 (URN)10.1007/978-3-031-82170-7_12 (DOI)2-s2.0-105012012683 (Scopus ID)
Note

Part of ISBN 9783031821691, 9783031821707

QC 20250820

Available from: 2025-08-20 Created: 2025-08-20 Last updated: 2025-08-20Bibliographically approved
Liu, Z., Liu, S., Wang, T., Wang, L. & Wang, X. V. (2025). Establishment and Synchronisation of Digital Twins for Multi-robot Systems in Manufacturing. In: 58th CIRP Conference on Manufacturing Systems, CMS 2025: . Paper presented at 58th CIRP Conference on Manufacturing Systems, CMS 2025, Twente, Netherlands, Kingdom of the, Apr 13 2025 - Apr 16 2025 (pp. 419-424). Elsevier BV
Open this publication in new window or tab >>Establishment and Synchronisation of Digital Twins for Multi-robot Systems in Manufacturing
Show others...
2025 (English)In: 58th CIRP Conference on Manufacturing Systems, CMS 2025, Elsevier BV , 2025, p. 419-424Conference paper, Published paper (Refereed)
Abstract [en]

In Industry 5.0, digital twins have emerged as powerful tools for revolutionizing the operation and control of industrial robots. However, a critical challenge is how to effectively synchronise the establishment and ongoing operations of physical devices with their virtual counterparts to ensure seamless performances. To address the challenge, this paper introduces a state machine-driven method to orchestrate hardware interface establishment and synchronisation processes for multi-robot systems in manufacturing. By leveraging state machines to model the lifecycle of hardware interfaces and their corresponding controllers, a systematic solution is provided for managing transitions and real-time synchronisation across multiple industrial robots. It not only enhances the initialisation efficiency but also ensures consistent system operations. The proposed method is validated through detailed case studies that demonstrate visible improvements in the deployment of manufacturing systems containing multiple industrial robots with different vendors and protocol interfaces. This work contributes to constructing digital twins that can dynamically adapt to evolving industrial environments.

Place, publisher, year, edition, pages
Elsevier BV, 2025
Keywords
Digital Twins, Multi-robot Systems in Manufacturing, Synchronisation, System Establishment
National Category
Computer Systems Robotics and automation Computer Sciences
Identifiers
urn:nbn:se:kth:diva-368826 (URN)10.1016/j.procir.2025.02.152 (DOI)2-s2.0-105009410507 (Scopus ID)
Conference
58th CIRP Conference on Manufacturing Systems, CMS 2025, Twente, Netherlands, Kingdom of the, Apr 13 2025 - Apr 16 2025
Note

QC 20250902

Available from: 2025-09-02 Created: 2025-09-02 Last updated: 2025-09-02Bibliographically approved
Zhao, S., Liu, S., Jiang, Y., Zhao, B., Lv, Y., Zhang, J., . . . Zhong, R. Y. (2025). Industrial Foundation Models (IFMs) for intelligent manufacturing: A systematic review. Journal of manufacturing systems, 82, 420-448
Open this publication in new window or tab >>Industrial Foundation Models (IFMs) for intelligent manufacturing: A systematic review
Show others...
2025 (English)In: Journal of manufacturing systems, ISSN 0278-6125, E-ISSN 1878-6642, Vol. 82, p. 420-448Article, review/survey (Refereed) Published
Abstract [en]

The remarkable success of Large Foundation Models (LFMs) has demonstrated their tremendous potential for manufacturing and sparked significant interest in the exploration of Industrial Foundation Models (IFMs). This study provides a comprehensive review of the current state of IFMs and their applications in intelligent manufacturing. It conducts an in-depth analysis from three perspectives, including data level, model level, and application level. The definition and framework of IFMs are discussed with a comparison to LFMs across these three perspectives. In addition, this paper provides a brief overview of the advancements in IFMs development across different countries, institutions, and regions. It explores the current application of IFMs, including Industrial Domain Models and Industrial Task Models, which are specifically designed for various industrial domains and tasks. Furthermore, key technologies critical to the training of IFMs are explored, such as data pre-processing, model fine-tuning, prompt engineering, and retrieval-augmented generation. This paper also highlights the essential capabilities of IFMs and their typical applications throughout the manufacturing lifecycle. Finally, it discusses the current challenges and outlines potential future research directions. This study aims to inspire new ideas for advancing IFMs and accelerating the evolution of intelligent manufacturing.

Place, publisher, year, edition, pages
Elsevier BV, 2025
Keywords
Industrial Foundation Models (IFMs), Intelligent manufacturing, Large Foundation Models (LFMs)
National Category
Production Engineering, Human Work Science and Ergonomics
Identifiers
urn:nbn:se:kth:diva-368893 (URN)10.1016/j.jmsy.2025.06.011 (DOI)001532863200001 ()2-s2.0-105009886814 (Scopus ID)
Note

QC 20250822

Available from: 2025-08-22 Created: 2025-08-22 Last updated: 2025-12-08Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-1909-0507

Search in DiVA

Show all publications