kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Alternative names
Publications (10 of 119) Show all publications
Zhu, X., Henningsson, J., Martensson, P., Hanson, L., Björkman, M. & Maki, A. (2026). Designing Synthetic Active Learning for model refinement in manufacturing parts detection. Journal of manufacturing systems, 84, 68-84
Open this publication in new window or tab >>Designing Synthetic Active Learning for model refinement in manufacturing parts detection
Show others...
2026 (English)In: Journal of manufacturing systems, ISSN 0278-6125, E-ISSN 1878-6642, Vol. 84, p. 68-84Article, review/survey (Refereed) Published
Abstract [en]

This paper introduces Synthetic Active Learning (SAL), a fully automatic model refinement strategy for manufacturing parts detection using only synthetic data actively generated with domain randomization for training. SAL iteratively updates the detection model by identifying its weaknesses, such as in specific categories, materials, or object sizes, using custom evaluators, and generating targeted synthetic data to address them; it selectively synthesizes new useful data with respect to active learning, where traditionally humans in the loop select data to label. During each iteration, model training and data generation occur simultaneously to improve efficiency. Evaluated on four use cases from two industrial datasets, SAL achieved mAP@50 improvements of 2 to 6% percentage points over static learning, which refers to training on a fixed, pre-generated dataset. It also showed notable gains in underperforming categories, leading to more balanced performance across classes. Another benefit is that it uses a consistent configuration across multiple use cases, avoiding the need for extensive hyperparameter tuning common in prior domain randomization studies. Given its encouraging performance across diverse scenarios, we believe that SAL can scale to broader industrial applications where training can be fully or mostly based on synthetic data.

Place, publisher, year, edition, pages
Elsevier BV, 2026
Keywords
Synthetic data, Domain randomization, Object detection, Active learning, Automation
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-377240 (URN)10.1016/j.jmsy.2025.11.023 (DOI)001636775700001 ()2-s2.0-105023671266 (Scopus ID)
Note

QC 20260224

Available from: 2026-02-24 Created: 2026-02-24 Last updated: 2026-02-24Bibliographically approved
Khanna, P., Rajabi, N., Kanik, S. U. e., Kragic Jensfelt, D., Björkman, M. & Smith, C. (2026). Early detection of human handover intentions in human–robot collaboration: Comparing EEG, gaze, and hand motion. Robotics and Autonomous Systems, 196, Article ID 105244.
Open this publication in new window or tab >>Early detection of human handover intentions in human–robot collaboration: Comparing EEG, gaze, and hand motion
Show others...
2026 (English)In: Robotics and Autonomous Systems, ISSN 0921-8890, E-ISSN 1872-793X, Vol. 196, article id 105244Article in journal (Refereed) Published
Abstract [en]

Human–robot collaboration (HRC) relies on accurate and timely recognition of human intentions to ensure seamless interactions. Among common HRC tasks, human-to-robot object handovers have been studied extensively for planning the robot's actions during object reception, assuming the human intention for object handover. However, distinguishing handover intentions from other actions has received limited attention. Most research on handovers has focused on visually detecting motion trajectories, which often results in delays or false detections when trajectories overlap. This paper investigates whether human intentions for object handovers are reflected in non-movement-based physiological signals. We conduct a multimodal analysis comparing three data modalities: electroencephalogram (EEG), gaze, and hand-motion signals. Our study aims to distinguish between handover-intended human motions and non-handover motions in an HRC setting, evaluating each modality's performance in predicting and classifying these actions before and after human movement initiation. We develop and evaluate human intention detectors based on these modalities, comparing their accuracy and timing in identifying handover intentions. To the best of our knowledge, this is the first study to systematically develop and test intention detectors across multiple modalities within the same experimental context of human–robot handovers. Our analysis reveals that handover intention can be detected from all three modalities. Nevertheless, gaze signals are the earliest as well as the most accurate to classify the motion as intended for handover or non-handover.

Place, publisher, year, edition, pages
Elsevier BV, 2026
Keywords
EEG, Gaze, Human–robot collaboration (HRC), Human–robot handovers, Motion analysis
National Category
Robotics and automation
Identifiers
urn:nbn:se:kth:diva-373139 (URN)10.1016/j.robot.2025.105244 (DOI)001619701400001 ()2-s2.0-105021346666 (Scopus ID)
Note

QC 20251121

Available from: 2025-11-21 Created: 2025-11-21 Last updated: 2026-05-29Bibliographically approved
Tarle, M., Larsson, M., Ingeström, G. & Björkman, M. (2026). Reinforcement Learning for Optimizing FACTS Setpoints With Limited Set of Measurements. IEEE Open Access Journal of Power and Energy, 13, 51-63
Open this publication in new window or tab >>Reinforcement Learning for Optimizing FACTS Setpoints With Limited Set of Measurements
2026 (English)In: IEEE Open Access Journal of Power and Energy, E-ISSN 2687-7910, Vol. 13, p. 51-63Article in journal (Refereed) Published
Abstract [en]

Coordinated control of Flexible AC Transmission Systems (FACTS) setpoints can significantly enhance power flow and voltage control. However, optimizing the setpoints of multiple FACTS devices in real-world systems remains uncommon, partly due to challenges in model-based control. Data-driven approaches, such as reinforcement learning (RL), offer a promising alternative for coordinated control. In this work, we address a setting where a useful real-time network model is unavailable. Recognizing the increasing deployment of Phasor Measurement Units (PMUs) for advanced monitoring and control, we consider having access to a few but reliable measurements and a constraint violation signal. Under these assumptions, we demonstrate on several scenarios on the IEEE 14-bus and IEEE 57-bus systems that an RL-based optimization of FACTS setpoints can substantially reduce voltage deviations compared to a fixed-setpoint baseline. To improve robustness and prevent unobserved constraint violations, we show that a complete, albeit simple, constraint violation signal is necessary. As an alternative to relying on such a signal, Dynamic Mode Decomposition is proposed to determine new PMU placements, thereby reducing the risk of unobserved constraint violations. Finally, to assess the gap to an optimal policy, we benchmark the RL-based agent against a model-based optimal controller with perfect information.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2026
Keywords
Decision support systems, Flexible AC Transmission Systems (FACTS), power system control, reinforcement learning
National Category
Control Engineering Other Electrical Engineering, Electronic Engineering, Information Engineering
Identifiers
urn:nbn:se:kth:diva-374972 (URN)10.1109/OAJPE.2025.3645591 (DOI)001666906400001 ()2-s2.0-105025685055 (Scopus ID)
Note

QC 20260127

Available from: 2026-01-09 Created: 2026-01-09 Last updated: 2026-05-29Bibliographically approved
Naoum, A., Khanna, P., Yadollahi, E., Björkman, M. & Smith, C. (2025). Adapting Robot's Explanation for Failures Based on Observed Human Behavior in Human-Robot Collaboration. In: IROS 2025 - 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, Conference Proceedings: . Paper presented at 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2025, Hangzhou, China, Oct 19 2025 - Oct 25 2025 (pp. 15087-15094). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Adapting Robot's Explanation for Failures Based on Observed Human Behavior in Human-Robot Collaboration
Show others...
2025 (English)In: IROS 2025 - 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, Conference Proceedings, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 15087-15094Conference paper, Published paper (Refereed)
Abstract [en]

This work aims to interpret human behavior to anticipate potential user confusion when a robot provides explanations for failure, allowing the robot to adapt its explanations for more natural and efficient collaboration. Using a dataset [1] that included facial emotion detection, eye gaze estimation, and gestures from 55 participants in a user study [2], we analyzed how human behavior changed in response to different types of failures and varying explanation levels. Our goal is to assess whether human collaborators are ready to accept less detailed explanations without inducing confusion. We formulate a data-driven predictor to predict human confusion during robot failure explanations. We also propose and evaluate a mechanism, based on the predictor, to adapt the explanation level according to observed human behavior. The promising results from this evaluation indicate the potential of this research in adapting a robot's explanations for failures to enhance the collaborative experience.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
National Category
Robotics and automation Computer Sciences Human Computer Interaction Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-377814 (URN)10.1109/IROS60139.2025.11245659 (DOI)2-s2.0-105029956695 (Scopus ID)
Conference
2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2025, Hangzhou, China, Oct 19 2025 - Oct 25 2025
Note

Part of ISBN 9798331543938

QC 20260310

Available from: 2026-03-10 Created: 2026-03-10 Last updated: 2026-03-10Bibliographically approved
Zhu, X., Mårtensson, P., Hanson, L., Björkman, M. & Maki, A. (2025). Automated assembly quality inspection by deep learning with 2D and 3D synthetic CAD data. Journal of Intelligent Manufacturing, 36(4), 2567-2582, Article ID e222.
Open this publication in new window or tab >>Automated assembly quality inspection by deep learning with 2D and 3D synthetic CAD data
Show others...
2025 (English)In: Journal of Intelligent Manufacturing, ISSN 0956-5515, E-ISSN 1572-8145, Vol. 36, no 4, p. 2567-2582, article id e222Article in journal (Refereed) Published
Abstract [en]

In the manufacturing industry, automatic quality inspections can lead to improved product quality and productivity. Deep learning-based computer vision technologies, with their superior performance in many applications, can be a possible solution for automatic quality inspections. However, collecting a large amount of annotated training data for deep learning is expensive and time-consuming, especially for processes involving various products and human activities such as assembly. To address this challenge, we propose a method for automated assembly quality inspection using synthetic data generated from computer-aided design (CAD) models. The method involves two steps: automatic data generation and model implementation. In the first step, we generate synthetic data in two formats: two-dimensional (2D) images and three-dimensional (3D) point clouds. In the second step, we apply different state-of-the-art deep learning approaches to the data for quality inspection, including unsupervised domain adaptation, i.e., a method of adapting models across different data distributions, and transfer learning, which transfers knowledge between related tasks. We evaluate the methods in a case study of pedal car front-wheel assembly quality inspection to identify the possible optimal approach for assembly quality inspection. Our results show that the method using Transfer Learning on 2D synthetic images achieves superior performance compared with others. Specifically, it attained 95% accuracy through fine-tuning with only five annotated real images per class. With promising results, our method may be suggested for other similar quality inspection use cases. By utilizing synthetic CAD data, our method reduces the need for manual data collection and annotation. Furthermore, our method performs well on test data with different backgrounds, making it suitable for different manufacturing environments.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
Assembly quality inspection, Computer vision, Point cloud, Synthetic data, Transfer learning, Unsupervised domain adaptation
National Category
Computer Sciences Production Engineering, Human Work Science and Ergonomics
Identifiers
urn:nbn:se:kth:diva-363099 (URN)10.1007/s10845-024-02375-6 (DOI)001205028300001 ()2-s2.0-105002924620 (Scopus ID)
Note

QC 20250506

Available from: 2025-05-06 Created: 2025-05-06 Last updated: 2025-11-12Bibliographically approved
Styrud, J., Iovino, M., Norrlöf, M., Björkman, M. & Smith, C. (2025). Automatic Behavior Tree Expansion with LLMs for Robotic Manipulation. In: 2025 IEEE International Conference on Robotics and Automation, ICRA 2025: . Paper presented at 2025 IEEE International Conference on Robotics and Automation, ICRA 2025, Atlanta, United States of America, May 19 2025 - May 23 2025 (pp. 1225-1232). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Automatic Behavior Tree Expansion with LLMs for Robotic Manipulation
Show others...
2025 (English)In: 2025 IEEE International Conference on Robotics and Automation, ICRA 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 1225-1232Conference paper, Published paper (Refereed)
Abstract [en]

Robotic systems for manipulation tasks are increasingly expected to be easy to configure for new tasks or unpredictable environments, while keeping a transparent policy that is readable and verifiable by humans. We propose the method BEhavior TRee eXPansion with Large Language Models (BETR-XP-LLM) to dynamically and automatically expand and configure Behavior Trees as policies for robot control. The method utilizes an LLM to resolve errors outside the task planner's capabilities, both during planning and execution. We show that the method is able to solve a variety of tasks and failures and permanently update the policy to handle similar problems in the future.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
National Category
Robotics and automation Computer Sciences Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-371379 (URN)10.1109/ICRA55743.2025.11127942 (DOI)001582497400110 ()2-s2.0-105016707385 (Scopus ID)
Conference
2025 IEEE International Conference on Robotics and Automation, ICRA 2025, Atlanta, United States of America, May 19 2025 - May 23 2025
Note

Part of ISBN 9798331541392

QC 20251010

Available from: 2025-10-10 Created: 2025-10-10 Last updated: 2026-05-29Bibliographically approved
Zhu, X., Henningsson, J., Li, D., Mårtensson, P., Hanson, L., Björkman, M. & Maki, A. (2025). Domain Randomization for Object Detection in Manufacturing Applications Using Synthetic Data: A Comprehensive Study. In: 2025 IEEE International Conference on Robotics and Automation, ICRA 2025: . Paper presented at 2025 IEEE International Conference on Robotics and Automation, ICRA 2025, Atlanta, United States of America, May 19 2025 - May 23 2025 (pp. 16715-16721). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Domain Randomization for Object Detection in Manufacturing Applications Using Synthetic Data: A Comprehensive Study
Show others...
2025 (English)In: 2025 IEEE International Conference on Robotics and Automation, ICRA 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 16715-16721Conference paper, Published paper (Refereed)
Abstract [en]

This paper addresses key aspects of domain randomization in generating synthetic data for manufacturing object detection applications. To this end, we present a comprehensive data generation pipeline that reflects different factors: object characteristics, background, illumination, camera settings, and post-processing. We also introduce the Synthetic Industrial Parts Object Detection dataset (SIP15-OD) consisting of 15 objects from three industrial use cases under varying environments as a test bed for the study, while also employing an industrial dataset publicly available for robotic applications. In our experiments, we present more abundant results and insights into the feasibility as well as challenges of sim-toreal object detection. In particular, we identified material properties, rendering methods, post-processing, and distractors as important factors. Our method, leveraging these, achieves top performance on the public dataset with Yolov8 models trained exclusively on synthetic data; mAP@50 scores of 96.4% for the robotics dataset, and 94.1%, 99.5%, and 95.3% across three of the SIP15-OD use cases, respectively. The results showcase the effectiveness of the proposed domain randomization, potentially covering the distribution close to real data for the applications.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
National Category
Computer graphics and computer vision Computer Sciences
Identifiers
urn:nbn:se:kth:diva-371386 (URN)10.1109/ICRA55743.2025.11128647 (DOI)001614889900500 ()2-s2.0-105016571384 (Scopus ID)
Conference
2025 IEEE International Conference on Robotics and Automation, ICRA 2025, Atlanta, United States of America, May 19 2025 - May 23 2025
Note

Part of ISBN 9798331541392

QC 20251009

Available from: 2025-10-09 Created: 2025-10-09 Last updated: 2026-05-29Bibliographically approved
Rajabi, N., Zanettin, I., Ribeiro, A. H., Vasco, M., Björkman, M., Lundström, J. N. & Kragic Jensfelt, D. (2025). Exploring the feasibility of olfactory brain–computer interfaces. Scientific Reports, 15(1)
Open this publication in new window or tab >>Exploring the feasibility of olfactory brain–computer interfaces
Show others...
2025 (English)In: Scientific Reports, E-ISSN 2045-2322, Vol. 15, no 1Article in journal (Refereed) Published
Abstract [en]

In this study, we explore the feasibility of single-trial predictions of odor registration in the brain using olfactory bio-signals. We focus on two main aspects: input data modality and the processing model. For the first time, we assess the predictability of odor registration from novel electrobulbogram (EBG) recordings, both in sensor and source space, and compare these with commonly used electroencephalogram (EEG) signals. Despite having fewer data channels, EBG shows comparable performance to EEG. We also examine whether breathing patterns contain relevant information for this task. By comparing a logistic regression classifier, which requires hand-crafted features, with an end-to-end convolutional deep neural network, we find that end-to-end approaches can be as effective as classic methods. However, due to the high dimensionality of the data, the current dataset is insufficient for either classifier to robustly differentiate odor and non-odor trials. Finally, we identify key challenges in olfactory BCIs and suggest future directions for improving odor detection systems.

Place, publisher, year, edition, pages
United Kingdom: Springer Nature, 2025
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-379335 (URN)10.1038/s41598-025-01488-z (DOI)001496076300032 ()40419502 (PubMedID)2-s2.0-105006408347 (Scopus ID)
Note

QC 20260416

Available from: 2026-04-16 Created: 2026-04-16 Last updated: 2026-04-16Bibliographically approved
Rajabi, N., Ribeiro, A. H., Vasco, M., Taleb, F., Björkman, M. & Kragic Jensfelt, D. (2025). Human-Aligned Image Models Improve Visual Decoding from the Brain. In: Proceedings of the 42nd International Conference on Machine Learning: . Paper presented at International Conference on Machine Learning, 13-19 July 2025, Vancouver, Canada. MLResearchPress
Open this publication in new window or tab >>Human-Aligned Image Models Improve Visual Decoding from the Brain
Show others...
2025 (English)In: Proceedings of the 42nd International Conference on Machine Learning, MLResearchPress , 2025Conference paper, Published paper (Refereed)
Abstract [en]

Decoding visual images from brain activity has significant potential for advancing brain-computer interaction and enhancing the understanding of human perception. Recent approaches align the representation spaces of images and brain activity to enable visual decoding. In this paper, we introduce the use of human-aligned image encoders to map brain signals to images. We hypothesize that these models more effectively capture perceptual attributes associated with the rapid visual stimuli presentations commonly used in visual brain data recording experiments. Our empirical results support this hypothesis, demonstrating that this simple modification improves image retrieval accuracy by up to 21% compared to state-of-the-art methods. Comprehensive experiments confirm consistent performance improvements across diverse EEG architectures, image encoders, alignment methods, participants, and brain imaging modalities.

Place, publisher, year, edition, pages
MLResearchPress, 2025
Series
Proceedings of Machine Learning Research, ISSN 2640-3498
Keywords
Computer Science: Computer Vision and Pattern Recognition, Computer Science: Machine Learning
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-379306 (URN)
Conference
International Conference on Machine Learning, 13-19 July 2025, Vancouver, Canada
Note

QC 20260416

Available from: 2026-04-16 Created: 2026-04-16 Last updated: 2026-04-16Bibliographically approved
Rajabi, N., Ribeiro, A. H., Vasco, M., Taleb, F., Björkman, M. & Kragic Jensfelt, D. (2025). Human-Aligned Image Models Improve Visual Decoding from the Brain. In: Singh, A Fazel, M Hsu, D Lacoste-Julien, S Berkenkamp, F Maharaj, T Wagstaff, K Zhu, J (Ed.), International Conference On Machine Learning: . Paper presented at 42nd International Conference on Machine Learning-ICML-Annual, JUL 13-19, 2025, Vancouver, CANADA (pp. 51009-51038). JMLR-JOURNAL MACHINE LEARNING RESEARCH, 267
Open this publication in new window or tab >>Human-Aligned Image Models Improve Visual Decoding from the Brain
Show others...
2025 (English)In: International Conference On Machine Learning / [ed] Singh, A Fazel, M Hsu, D Lacoste-Julien, S Berkenkamp, F Maharaj, T Wagstaff, K Zhu, J, JMLR-JOURNAL MACHINE LEARNING RESEARCH , 2025, Vol. 267, p. 51009-51038Conference paper, Published paper (Refereed)
Abstract [en]

Decoding visual images from brain activity has significant potential for advancing brain-computer interaction and enhancing the understanding of human perception. Recent approaches align the representation spaces of images and brain activity to enable visual decoding. In this paper, we introduce the use of human-aligned image encoders to map brain signals to images. We hypothesize that these models more effectively capture perceptual attributes associated with the rapid visual stimuli presentations commonly used in visual brain data recording experiments. Our empirical results support this hypothesis, demonstrating that this simple modification improves image retrieval accuracy by up to 21% compared to state-of-the-art methods. Comprehensive experiments confirm consistent performance improvements across diverse EEG architectures, image encoders, alignment methods, participants, and brain imaging modalities(1).

Place, publisher, year, edition, pages
JMLR-JOURNAL MACHINE LEARNING RESEARCH, 2025
Series
Proceedings of Machine Learning Research, ISSN 2640-3498
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-382257 (URN)001693162000235 ()2-s2.0-105023639824 (Scopus ID)
Conference
42nd International Conference on Machine Learning-ICML-Annual, JUL 13-19, 2025, Vancouver, CANADA
Note

QC 20260527

Available from: 2026-05-27 Created: 2026-05-27 Last updated: 2026-05-27Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0003-0579-3372

Search in DiVA

Show all publications