kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Alternative names
Publications (7 of 7) Show all publications
Squires, L., Giraldez Chavez, J. H., Nilsson, A., Käll, L. & Payne, S. H. (2026). Better Inputs, Better Learning: A Peptide Embedding Tutorial for Proteomic Mass Spectrometry. Journal of Proteome Research, 25(2), 1160-1165
Open this publication in new window or tab >>Better Inputs, Better Learning: A Peptide Embedding Tutorial for Proteomic Mass Spectrometry
Show others...
2026 (English)In: Journal of Proteome Research, ISSN 1535-3893, E-ISSN 1535-3907, Vol. 25, no 2, p. 1160-1165Article in journal (Refereed) Published
Abstract [en]

Mass spectrometry proteomics creates complex data representing the peptide/protein contents of biological samples. Various types of machine learning have been central to computational methods used to identify peptides from tandem mass spectra and numerous other aspects of the data analysis process. As deep learning has emerged as a powerful machine learning method for modeling and interpreting data, computational proteomics researchers have leveraged large publicly available data sets to train machine learning models to predict peptide fragmentation spectra and liquid chromatography retention time. Resources like proteomicsML offer extensive demonstrative tutorials for these learning tasks and are closing the gap between the proteomics and machine learning communities. However, in these and other educational materials on deep learning, the critical step of preparing data for learning is frequently omitted. Prior to learning, peptide strings must be converted into a numeric format─an embedding. There are many different peptide embeddings, and some vastly outperform others. Yet the process for creating an embedding, and also the rationale for choosing a specific embedding, is rarely discussed in our proteomics literature. In this technical note, we introduce four Google Colab notebooks to teach peptide embeddings. The series walks users through five different peptide-embedding strategies─ from simplistic single-number encodings to state-of-the-art pretrained embeddings─ through both code examples and narrative descriptions. The final notebook compares the five embeddings in a head-to-head benchmark. By making these notebooks free, we hope to lower the barrier for researchers who want to bring modern deep learning into their proteomics workflows.

Place, publisher, year, edition, pages
American Chemical Society (ACS), 2026
Keywords
embedding, encoding, machine learning, peptide, proteomics AI, proteomics education, tutorials
National Category
Bioinformatics (Computational Biology) Bioinformatics and Computational Biology
Identifiers
urn:nbn:se:kth:diva-377327 (URN)10.1021/acs.jproteome.5c00563 (DOI)001661059700001 ()41528974 (PubMedID)2-s2.0-105029493700 (Scopus ID)
Note

QC 20260227

Available from: 2026-02-27 Created: 2026-02-27 Last updated: 2026-02-27Bibliographically approved
Nilsson, A. (2026). Learning Representations for Tandem Mass Spectra: Self-Supervised Methods and Inductive Biases. (Licentiate dissertation). Stockholm: KTH Royal Institute of Technology
Open this publication in new window or tab >>Learning Representations for Tandem Mass Spectra: Self-Supervised Methods and Inductive Biases
2026 (English)Licentiate thesis, comprehensive summary (Other academic)
Abstract [en]

Mass spectrometry (MS) is central to modern proteomics, enabling analysis of proteins and peptides based on their mass-to-charge ratio. Tandem mass spectrometry (MS2) encodes peptide fragmentation patterns and forms the basis for sequence identification. While database search has long dominated this process, deep learning has opened new paths for the direct interpretation of spectra. This thesis investigates how neural networks can learn representations of MS2 spectra. Two complementary research directions are explored.

First, selected self-supervised pretraining strategies are evaluated through controlled downstream experiments using encoders pretrained on unlabeled MS2 corpora. Self-distillation yields global embeddings that implicitly encode aspects of peptide chemical properties, and masked autoencoding provides modest improvements in de novo optimization and accuracy. However, the resulting improvements fall short of state-of-the-art supervised de novo sequencing performance.

Second, we introduce Pairwise Attention, a transformer architecture that incorporates a domain-aligned relational inductive bias by conditioning attention on pairwise mass differences between peaks. This yields consistent performance improvements on standard de novo sequencing benchmarks and strong generalization across datasets.

Overall, the results show that self-supervised learning can recover meaningful structure from raw MS2 data, while architectural inductive biases currently offer the most robust and reliable gains for de novo peptide sequencing.

Abstract [sv]

Masspektrometrin (MS) är central inom modern proteomik och möjliggör analysav proteiner och peptider baserat på deras massa. Tandem-masspektrometri (MS2)kodar fragmenteringsmönster för peptider och utgör grunden för sekvensidentifiering. Även om databassökning länge har dominerat denna process har djupinlärning öppnat nya möjligheter för direkt tolkning av spektra.

Denna avhandling undersöker hur neurala nätverk kan lära sig representationer av MS2-spektra. Två kompletterande forskningsinriktningar studeras.

Först utvärderas utvalda självövervakade förträningsstrategier genom kontrollerade experiment med encoders som förtränats på oetiketterade MS2-korpusar. Självdistillation ger globala inbäddningar som implicit kodar aspekter av peptiders kemiska egenskaper, och masked autoencoding ger måttliga förbättringar i de novo-precision. De resulterande förbättringarna når dock inte upp till prestandan hos dagens state-of-the-art-metoder för övervakad de novo-sekvensering.

Sedan introduceras Pairwise Attention, en transformerarkitektur som inkorporerar en domänanpassad induktiv bias genom att villkora Attention på parvisa masskillnader mellan toppar. Detta ger prestandaförbättringar på etablerade de novo-sekvenseringsbenchmarkar samt stark generalisering över dataset.

Sammantaget visar resultaten att självövervakad inlärning kan återvinna meningsfull struktur ur råa MS2-data, medan induktiva biaser för närvarande erbjuder de mest robusta förbättringarna för de novo-peptidsekvensering.

Place, publisher, year, edition, pages
Stockholm: KTH Royal Institute of Technology, 2026. p. 45
Series
TRITA-CBH-FOU ; 2026:21
Keywords
Mass Spectrometry, Deep Learning, De Novo Sequencing, Self-Supervised Learning
National Category
Bioinformatics and Computational Biology
Research subject
Biotechnology
Identifiers
urn:nbn:se:kth:diva-378805 (URN)978-91-8106-586-2 (ISBN)
Presentation
2026-04-17, Pascal, Gamma-6, Tomtebodavägen 23, Solna, Stockholm, 13:15 (English)
Opponent
Supervisors
Note

QC 2026-03-27

Available from: 2026-03-27 Created: 2026-03-27 Last updated: 2026-03-30Bibliographically approved
Lapin, J., Nilsson, A., Wilhelm, M. & Käll, L. (2025). Pairwise Attention: Leveraging Mass Differences to Enhance De Novo Sequencing of Mass Spectra. Journal of Proteome Research, 24(7), 3722-3730
Open this publication in new window or tab >>Pairwise Attention: Leveraging Mass Differences to Enhance De Novo Sequencing of Mass Spectra
2025 (English)In: Journal of Proteome Research, ISSN 1535-3893, E-ISSN 1535-3907, Vol. 24, no 7, p. 3722-3730Article in journal (Refereed) Published
Abstract [en]

A fundamental challenge in mass spectrometry-based proteomics is determining which peptide generated a given MS2 spectrum. Peptide sequencing typically relies on matching spectra against a known sequence database, which in some applications is not available. Deep learning-based de novo sequencing can address this limitation by directly predicting peptide sequences from MS2 data. We have seen the application of the transformer architecture to de novo sequencing produce state-of-the-art results on the so-called nine-species benchmark. In this study, we propose an improved transformer encoder inspired by the heuristics used in the manual interpretation of spectra. We modify the attention mechanism with a learned bias based on pairwise mass differences, termed Pairwise Attention (PA). Adding PA improves average peptide precision at 100% coverage by 12.7% (5.9 percentage points) over our base transformer on the original nine-species benchmark. We have also achieved a 7.4% increase over the previously published model Casanovo. Our MS2 encoding strategy is largely orthogonal to other transformer-based models encoding MS2 spectra, enabling straightforward integration into existing deep-learning approaches. Our results show that integrating domain-specific knowledge into transformers boosts de novo sequencing performance.

Place, publisher, year, edition, pages
American Chemical Society (ACS), 2025
Keywords
Attention, De novo sequencing, Mass spectrometry, MS2, Proteomics, Transformers
National Category
Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:kth:diva-364463 (URN)10.1021/acs.jproteome.5c00063 (DOI)001500093400001 ()40454436 (PubMedID)2-s2.0-105007317018 (Scopus ID)
Funder
Knut and Alice Wallenberg Foundation, KAW 2022.0032
Note

QC 20260127

Available from: 2025-06-12 Created: 2025-06-12 Last updated: 2026-05-12Bibliographically approved
Nilsson, A., Wijk, K., Gutha, S. b., Englesson, E., Hotti, A., Saccardi, C., . . . Azizpour, H. (2024). Indirectly Parameterized Concrete Autoencoders. In: International Conference on Machine Learning, ICML 2024: . Paper presented at 41st International Conference on Machine Learning, ICML 2024, Vienna, Austria, Jul 21 2024 - Jul 27 2024 (pp. 38237-38252). ML Research Press
Open this publication in new window or tab >>Indirectly Parameterized Concrete Autoencoders
Show others...
2024 (English)In: International Conference on Machine Learning, ICML 2024, ML Research Press , 2024, p. 38237-38252Conference paper, Published paper (Refereed)
Abstract [en]

Feature selection is a crucial task in settings where data is high-dimensional or acquiring the full set of features is costly. Recent developments in neural network-based embedded feature selection show promising results across a wide range of applications. Concrete Autoencoders (CAEs), considered state-of-the-art in embedded feature selection, may struggle to achieve stable joint optimization, hurting their training time and generalization. In this work, we identify that this instability is correlated with the CAE learning duplicate selections. To remedy this, we propose a simple and effective improvement: Indirectly Parameterized CAEs (IP-CAEs). IP-CAEs learn an embedding and a mapping from it to the Gumbel-Softmax distributions' parameters. Despite being simple to implement, IP-CAE exhibits significant and consistent improvements over CAE in both generalization and training time across several datasets for reconstruction and classification. Unlike CAE, IP-CAE effectively leverages non-linear relationships and does not require retraining the jointly optimized decoder. Furthermore, our approach is, in principle, generalizable to Gumbel-Softmax distributions beyond feature selection.

Place, publisher, year, edition, pages
ML Research Press, 2024
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-353956 (URN)2-s2.0-85203808876 (Scopus ID)
Conference
41st International Conference on Machine Learning, ICML 2024, Vienna, Austria, Jul 21 2024 - Jul 27 2024
Note

QC 20240926

Available from: 2024-09-25 Created: 2024-09-25 Last updated: 2024-09-26Bibliographically approved
Nilsson, A. & Azizpour, H. (2024). Regularizing and Interpreting Vision Transformers by Patch Selection on Echocardiography Data. In: Proceedings of the 5th Conference on Health, Inference, and Learning, CHIL 2024: . Paper presented at 5th Annual Conference on Health, Inference, and Learning, CHIL 2024, New York, United States of America, Jun 27 2024 - Jun 28 2024 (pp. 155-168). ML Research Press
Open this publication in new window or tab >>Regularizing and Interpreting Vision Transformers by Patch Selection on Echocardiography Data
2024 (English)In: Proceedings of the 5th Conference on Health, Inference, and Learning, CHIL 2024, ML Research Press , 2024, p. 155-168Conference paper, Published paper (Refereed)
Abstract [en]

This work introduces a novel approach to model regularization and explanation in Vision Transformers (ViTs), particularly beneficial for small-scale but high-dimensional data regimes, such as in healthcare. We introduce stochastic embedded feature selection in the context of echocardiography video analysis, specifically focusing on the EchoNet-Dynamic dataset for the prediction of Left Ventricular Ejection Fraction (LVEF). Our proposed method, termed Gumbel Video Vision-Transformers (G-ViTs), augments Video Vision-Transformers (V-ViTs), a performant transformer architecture for videos with Concrete Autoencoders (CAEs), a common dataset-level feature selection technique, to enhance V-ViT’s generalization and interpretability. The key contribution lies in the incorporation of stochastic token selection individually for each video frame during training. Such token selection regularizes the training of V-ViT, improves its interpretability, and is achieved by differentiable sampling of categoricals using the Gumbel-Softmax distribution. Our experiments on EchoNet-Dynamic demonstrate a consistent and notable regularization effect. The G-ViT model outperforms both a random selection baseline and standard V-ViT. The G-ViT is also compared against recent works on EchoNet-Dynamic where it exhibits state-of-the-art performance among end-to-end learned methods. Finally, we explore model ex-plainability by visualizing selected patches, providing insights into how the G-ViT utilizes regions known to be crucial for LVEF prediction for humans. This proposed approach, therefore, extends beyond regularization, offering enhanced interpretability for ViTs.

Place, publisher, year, edition, pages
ML Research Press, 2024
National Category
Computer graphics and computer vision Computer Sciences
Identifiers
urn:nbn:se:kth:diva-353944 (URN)2-s2.0-85203788338 (Scopus ID)
Conference
5th Annual Conference on Health, Inference, and Learning, CHIL 2024, New York, United States of America, Jun 27 2024 - Jun 28 2024
Note

QC 20240926

Available from: 2024-09-25 Created: 2024-09-25 Last updated: 2025-02-01Bibliographically approved
Nilsson, A. & Azizpour, H. (2024). Regularizing and Interpreting Vision Transformers by Patch Selection on Echocardiography Data. In: Pollard, T Choi, E Singhal, P Hughes, M Sizikova, E Mortazavi, B Chen, I Wang, F Sarker, T McDermott, M Ghassemi, M (Ed.), CONFERENCE ON HEALTH, INFERENCE, AND LEARNING: . Paper presented at 5th Annual Conference on Health, Inference, and Learning (CHIL), JUN 27-28, 2024, New York, NY (pp. 155-168). The Journal of Machine Learning Research (JMLR), 248
Open this publication in new window or tab >>Regularizing and Interpreting Vision Transformers by Patch Selection on Echocardiography Data
2024 (English)In: CONFERENCE ON HEALTH, INFERENCE, AND LEARNING / [ed] Pollard, T Choi, E Singhal, P Hughes, M Sizikova, E Mortazavi, B Chen, I Wang, F Sarker, T McDermott, M Ghassemi, M, The Journal of Machine Learning Research (JMLR) , 2024, Vol. 248, p. 155-168Conference paper, Published paper (Refereed)
Abstract [en]

This work introduces a novel approach to model regularization and explanation in Vision Transformers (ViTs), particularly beneficial for small-scale but high-dimensional data regimes, such as in healthcare. We introduce stochastic embedded feature selection in the context of echocardiography video analysis, specifically focusing on the EchoNet-Dynamic dataset for the prediction of Left Ventricular Ejection Fraction (LVEF). Our proposed method, termed Gumbel Video Vision-Transformers (G-ViTs), augments Video Vision-Transformers (V-ViTs), a performant transformer architecture for videos with Concrete Autoencoders (CAEs), a common dataset-level feature selection technique, to enhance V-ViT's generalization and interpretability. The key contribution lies in the incorporation of stochastic token selection individually for each video frame during training. Such token selection regularizes the training of V-ViT, improves its interpretability, and is achieved by differentiable sampling of categoricals using the Gumbel-Softmax distribution. Our experiments on EchoNet-Dynamic demonstrate a consistent and notable regularization effect. The G-ViT model outperforms both a random selection baseline and standard V-ViT. The G-ViT is also compared against recent works on EchoNet-Dynamic where it exhibits state-of-the-art performance among end-to-end learned methods. Finally, we explore model explainability by visualizing selected patches, providing insights into how the G-ViT utilizes regions known to be crucial for LVEF prediction for humans. This proposed approach, therefore, extends beyond regularization, offering enhanced interpretability for ViTs.

Place, publisher, year, edition, pages
The Journal of Machine Learning Research (JMLR), 2024
Series
Proceedings of Machine Learning Research, ISSN 2640-3498
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-357514 (URN)001347132700011 ()
Conference
5th Annual Conference on Health, Inference, and Learning (CHIL), JUN 27-28, 2024, New York, NY
Note

QC 20241209

Available from: 2024-12-09 Created: 2024-12-09 Last updated: 2025-02-07Bibliographically approved
Nilsson, A. & Käll, L.Self-Supervised Learning for Tandem Mass Spectra: Methods, Dynamics and Downstream Effects.
Open this publication in new window or tab >>Self-Supervised Learning for Tandem Mass Spectra: Methods, Dynamics and Downstream Effects
(English)Manuscript (preprint) (Other academic)
Abstract [en]

Self-supervised learning provides a way to extract structure from tandem mass spectra without relying on peptide labels, which are typically obtained from database-search pipelines and therefore inherit their assumptions and biases. This work evaluates several self-supervised objectives: masked spectrum modeling, masked autoencoding, trinary m/zperturbation, and DINO-style self-distillation, using a shared transformer encoder trained on large unlabeled MS2 corpora. We analyze their training behavior, document characteristic failure modes such as collapse in DINO, and measure their effect on downstream de novo peptide sequencing and auxiliary prediction tasks.

Across controlled fine-tuning settings, masked autoencoding yields the most consistent improvements in de novo accuracy, with measurable gains even after very limited pretraining. DINO provides modest but reproducible improvements over scratch for de novo decoding and strong gains on global tasks, whereas the trinary perturbation objective produces only small and often inconsistent benefits. These results demonstrate that unsupervised objectives can recover meaningful structure from raw spectra, although the absolute de novo accuracies achieved here lie below those of state-of-the-art supervised systems, meaning the observed gains should be interpreted primarily as an initialization ablation rather than an indication of absolute model capability.

Overall, the study shows that self-supervision can influence MS2 representations in useful ways, clarifies which objectives are effective for current transformer architectures, and highlights the need for MS-specific pretraining tasks that more directly support high-quality sequence reconstruction

Keywords
Deep Learning, Mass Spectrometry, Self-Supervised Learning
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:kth:diva-378799 (URN)
Note

QC 20260331

Available from: 2026-03-27 Created: 2026-03-27 Last updated: 2026-03-31Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-3181-3800

Search in DiVA

Show all publications