kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Beam search decoder for enhancing sequence decoding speed in single-molecule peptide sequencing data
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Information Science and Engineering.ORCID iD: 0000-0002-6753-8548
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Information Science and Engineering.ORCID iD: 0000-0001-6630-243X
2023 (English)In: PloS Computational Biology, ISSN 1553-734X, E-ISSN 1553-7358, Vol. 19, no 11, article id e1011345Article in journal (Refereed) Published
Abstract [en]

Next-generation single-molecule protein sequencing technologies have the potential to significantly accelerate biomedical research. These technologies offer sensitivity and scalability for proteomic analysis. One auspicious method is fluorosequencing, which involves: cutting naturalized proteins into peptides, attaching fluorophores to specific amino acids, and observing variations in light intensity as one amino acid is removed at a time. The original peptide is classified from the sequence of light-intensity reads, and proteins can subsequently be recognized with this information. The amino acid step removal is achieved by attaching the peptides to a wall on the C-terminal and using a process called Edman Degradation to remove an amino acid from the N-Terminal. Even though a framework (Whatprot) has been proposed for the peptide classification task, processing times remain restrictive due to the massively parallel data acquisicion system. In this paper, we propose a new beam search decoder with a novel state formulation that obtains considerably lower processing times at the expense of only a slight accuracy drop compared to Whatprot. Furthermore, we explore how our novel state formulation may lead to even faster decoders in the future.

Place, publisher, year, edition, pages
Public Library of Science (PLoS) , 2023. Vol. 19, no 11, article id e1011345
National Category
Biochemistry Molecular Biology
Identifiers
URN: urn:nbn:se:kth:diva-340116DOI: 10.1371/journal.pcbi.1011345ISI: 001101898700004PubMedID: 37934778Scopus ID: 2-s2.0-85176315601OAI: oai:DiVA.org:kth-340116DiVA, id: diva2:1815184
Note

QC 20231128

Available from: 2023-11-28 Created: 2023-11-28 Last updated: 2025-12-05Bibliographically approved
In thesis
1. Algorithms and machine learning for single-molecule protein sequencing methods
Open this publication in new window or tab >>Algorithms and machine learning for single-molecule protein sequencing methods
2025 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Single-molecule protein sequencing (SMPS) technologies are powerfulalternatives to mass spectrometry, offering new opportunities for highresolutionproteomics. These technologies, including nanopores, nanogaps,and fluorosequencing, enable the direct identification of protein moleculesat single-molecule resolution. Their potential spans diverse applications,from supporting cutting-edge biological research to developing diagnosticsand therapeutics. However, SMPS platforms generate complex and noisysignals in large volumes, making computational analysis a key bottleneck inunlocking their full potential.This thesis addresses that challenge by developing scalable, modelinformedand data-driven algorithms tailored to SMPS data. Drawing ontools from statistical signal processing and machine learning, the work focuseson computational methods that improve signal denoising, inference accuracy,and runtime efficiency across several SMPS technologies.The contributions span three major sensing platforms. For nanogap tunnelingdevices, a fast and robust denoising algorithm is introduced to managethe heavy-tailed noise characteristic of electronic tunneling signals. Fornanopore DNA sensing, a physics-inspired data augmentation method is proposedto improve the generalization of neural networks without requiring additionalexperimental data. Alongside this data augmentation, the thesisintroduces a novel neural network architecture that leverages the augmentation’sbenefits and incorporates modern design principles, such as residualconnections and attention mechanisms, to outperform state-of-the-art modelson a nanopore classification task.Finally, for fluorosequencing, this thesis presents two complementary contributions:(i) a fast beam search decoder for peptide inference and (ii) anexpectation-maximization framework for protein abundance estimation. Theproposed decoder achieves up to a tenfold speedup over existing methods withonly minimal loss in accuracy. Building on its output, the EM-based proteininference framework enables efficient estimation of protein abundances frompeptide-level posteriors. We demonstrate that this approach not only improvesquantification accuracy on small-scale datasets but also scales to thefull human proteome with tractable computation times, offering a viable routetoward single-molecule proteomics at large scale. Together, these tools contributeto the broader effort of making SMPS computationally tractable atthe scale required for full-proteome and single-cell analyses.All methods in this thesis have been made available as open-source software,reflecting a commitment to reproducibility and to supporting the growingSMPS research community. Through the integration of domain knowledge,algorithmic design, and computational efficiency, this thesis aims topush the boundaries of what is achievable in next-generation proteomics.

Abstract [sv]

Singelmolekylär proteinsekvensering (SMPS) utgör ett kraftfullt komplement och alternativ till masspektrometri och öppnar för nya möjligheter inom högupplöst proteomik. Tekniker som nanoporer, nanogap-strukturer och fluorosekvensering möjliggör direkt identifiering av enskilda proteinmolekyler med singelmolekylupplösning. Användningsområdet är brett—från stöd för frontlinjens biologiska forskning till utveckling av diagnostik och terapier. Samtidigt genererar SMPS-plattformar komplexa och brusiga signaler i stora volymer, vilket gör den beräkningsmässiga analysen till ett centralt hinder för  att realisera teknikernas fulla potential.

Avhandlingen adresserar denna utmaning genom att utveckla skalbara, modellunderbyggda och datadrivna algoritmer specifikt anpassade för SMPS-data. Med utgångspunkt i statistisk signalbehandling och maskininlärning utvecklas metoder som förbättrar brusreducering, inferensnoggrannhet och beräkningseffektivitet över flera SMPS-tekniker.

Bidragen spänner över tre huvudplattformar. För nanogap-baserad tunneleringssensorik presenteras en snabb och robust algoritm för brusreducering som effektivt hanterar det tungsvansade brus som är typiskt för elektroniska tunneleringssignaler. För nanoporsbaserad DNA-avläsning introduceras en fysikinspirerad dataaugmentering som höjer neurala nätverks generaliseringsförmåga utan krav på ytterligare experimentella data. I anslutning därtill föreslås en ny neuronnätsarkitektur som drar nytta av augmenteringen och införlivar moderna designprinciper, bland annat residualkopplingar och uppmärksamhetsmekanismer, vilket sammantaget överträffar state-of-the-art avancerade metoder  på en nanoporklassificeringsuppgift.

För fluorosekvensering presenteras två kompletterande komponenter: (i) en snabb beam search-avkodare för peptid-inferens och (ii) ett ramverk för proteinkvantifiering baserat på Expectation Maximization (EM). Avkodaren är upp till tio gånger snabbare än befintliga metoder med endast marginell försämring i noggrannhet. Baserat på dess utdata möjliggör det EM-baserade proteininferensramverket effektiv skattning av proteinabundanser från posteriorer på peptidnivå. Vi visar att angreppssättet inte bara förbättrar kvantifieringsnoggrannheten på småskaliga dataset, utan även skalar till hela det mänskliga proteomet med hanterbara beräkningstider, och därmed erbjuder en praktiskt genomförbar väg mot singelmolekylär proteomik i stor skala. Tillsammans bidrar dessa verktyg till att göra SMPS beräkningsmässigt hanterligt i den skala som krävs för helproteom- och enkelcellsanalyser.

Samtliga metoder i avhandlingen har gjorts tillgängliga som programvara med öppen källkod, i linje med ett starkt åtagande för reproducerbarhet och för att stödja det växande forskningsfältet kring SMPS. Genom att förena domänkunskap, välgrundad algoritmdesign och beräkningseffektivitet syftar avhandlingen till att flytta fram gränserna för vad som är möjligt inom nästa generations proteomik.

Place, publisher, year, edition, pages
Kungliga Tekniska högskolan, 2025. p. 135
Series
TRITA-EECS-AVL ; 2025:86
Keywords
Signal processing, Hidden Markov Models, Expectation Maximization, CUSUM, Data augmentation, Convolutional Neural Networks
National Category
Other Electrical Engineering, Electronic Engineering, Information Engineering Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:kth:diva-370661 (URN)978-91-8106-409-4 (ISBN)
Public defence
2025-11-07, F3, Lindstedtvägen 26, Stockholm, 13:00 (English)
Opponent
Supervisors
Note

QC 20250930

Available from: 2025-09-30 Created: 2025-09-29 Last updated: 2025-10-14Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textPubMedScopus

Authority records

Kipen, JavierJaldén, Joakim

Search in DiVA

By author/editor
Kipen, JavierJaldén, Joakim
By organisation
Information Science and Engineering
In the same journal
PloS Computational Biology
BiochemistryMolecular Biology

Search outside of DiVA

GoogleGoogle Scholar

doi
pubmed
urn-nbn

Altmetric score

doi
pubmed
urn-nbn
Total: 148 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf