kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Perception-decision-execution coordination mechanism driven dynamic autonomous collaboration method for human-like collaborative robot based on multimodal large language model
Beijing Inst Technol, Sch Mech Engn, Beijing 100081, Peoples R China.
Beijing Inst Technol, Sch Mech Engn, Beijing 100081, Peoples R China; Minist Ind & Informat Technol, Beijing Inst Technol, Key Lab Ind Knowledge & Data Fus Technol & Applica, Beijing 100081, Peoples R China.
Beijing Inst Technol, Sch Mech Engn, Beijing 100081, Peoples R China.
Beijing Inst Technol, Sch Mech Engn, Beijing 100081, Peoples R China.
Show others and affiliations
2026 (English)In: Robotics and Computer-Integrated Manufacturing, ISSN 0736-5845, E-ISSN 1879-2537, Vol. 98, article id 103167Article in journal (Refereed) Published
Abstract [en]

With the advent of Industry 5.0, human-centric smart manufacturing is becoming a new paradigm for industrial transformation. Human-robot collaboration (HRC) is the hot topic of human-centric smart manufacturing. The emergence of large language model (LLM) provides significant opportunity for collaborative robot to promote the autonomous collaboration ability, which brings HRC into new era driven by embodied intelligence and more powerful robot. Therefore, a dynamic autonomous collaboration method inspired from looking-thinking-doing chain of human operators is proposed for human-like collaborative robot (HLCobot) in human-centric smart manufacturing based on multimodal large language model (MLLM), where perception-decision-execution coordination mechanism is constructed to appropriately distribute the abilities of MLLM in the dynamic operation chain of HRC. Firstly, a brain-inspired architecture with the integration of perception hub, decision hub, and execution hub is designed for dynamic autonomous collaboration. Secondly, the abilities of perception, decision, execution of HLCobot are realized by integrating MLLM, where the HLCobot can actively recognize the dynamic changes of HRC scenario by mimicking human operator and execute the correct motions to complete the necessary collaborative task autonomously. Additionally, a coordination mechanism among the agents of perception, decision, and execution is put forward to proceed the collaborative task smoothly. Finally, a case study of engine assembly is provided to demonstrate the effectiveness of the proposed method.

Place, publisher, year, edition, pages
Elsevier BV , 2026. Vol. 98, article id 103167
Keywords [en]
Human-centric smart manufacturing, Human-robot collaboration, Human-like collaborative robot, Dynamic autonomous collaboration, Multimodal large language model, Perception-decision-execution coordination
National Category
Robotics and automation
Identifiers
URN: urn:nbn:se:kth:diva-374796DOI: 10.1016/j.rcim.2025.103167ISI: 001595174800001Scopus ID: 2-s2.0-105018169281OAI: oai:DiVA.org:kth-374796DiVA, id: diva2:2027768
Note

QC 20260113

Available from: 2026-01-13 Created: 2026-01-13 Last updated: 2026-01-13Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Wang, Lihui

Search in DiVA

By author/editor
Wang, Lihui
By organisation
Production Engineering
In the same journal
Robotics and Computer-Integrated Manufacturing
Robotics and automation

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 42 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf