Better Inputs, Better Learning: A Peptide Embedding Tutorial for Proteomic Mass SpectrometryShow others and affiliations
2026 (English)In: Journal of Proteome Research, ISSN 1535-3893, E-ISSN 1535-3907, Vol. 25, no 2, p. 1160-1165Article in journal (Refereed) Published
Abstract [en]
Mass spectrometry proteomics creates complex data representing the peptide/protein contents of biological samples. Various types of machine learning have been central to computational methods used to identify peptides from tandem mass spectra and numerous other aspects of the data analysis process. As deep learning has emerged as a powerful machine learning method for modeling and interpreting data, computational proteomics researchers have leveraged large publicly available data sets to train machine learning models to predict peptide fragmentation spectra and liquid chromatography retention time. Resources like proteomicsML offer extensive demonstrative tutorials for these learning tasks and are closing the gap between the proteomics and machine learning communities. However, in these and other educational materials on deep learning, the critical step of preparing data for learning is frequently omitted. Prior to learning, peptide strings must be converted into a numeric format─an embedding. There are many different peptide embeddings, and some vastly outperform others. Yet the process for creating an embedding, and also the rationale for choosing a specific embedding, is rarely discussed in our proteomics literature. In this technical note, we introduce four Google Colab notebooks to teach peptide embeddings. The series walks users through five different peptide-embedding strategies─ from simplistic single-number encodings to state-of-the-art pretrained embeddings─ through both code examples and narrative descriptions. The final notebook compares the five embeddings in a head-to-head benchmark. By making these notebooks free, we hope to lower the barrier for researchers who want to bring modern deep learning into their proteomics workflows.
Place, publisher, year, edition, pages
American Chemical Society (ACS) , 2026. Vol. 25, no 2, p. 1160-1165
Keywords [en]
embedding, encoding, machine learning, peptide, proteomics AI, proteomics education, tutorials
National Category
Bioinformatics (Computational Biology) Bioinformatics and Computational Biology
Identifiers
URN: urn:nbn:se:kth:diva-377327DOI: 10.1021/acs.jproteome.5c00563ISI: 001661059700001PubMedID: 41528974Scopus ID: 2-s2.0-105029493700OAI: oai:DiVA.org:kth-377327DiVA, id: diva2:2042173
Note
QC 20260227
2026-02-272026-02-272026-02-27Bibliographically approved