kth.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Algorithms and Models in Nanopore DNA Sequencing: Advanced Decoding and Modeling with Hierarchical Hidden Markov Models
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Information Science and Engineering. (Division of Information Science and Engineering)ORCID iD: 0000-0003-1850-0946
2024 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Within less than four decades, the nanopore sequencing technology accelerated from an implausible idea scribbled on a notebook page to one of the decisive contributors to the complete sequence of the human genome. Its rapid evolution, particularly in recent years, is driven not only by its inherent innovation of the nanopores but also by synergistic advancements in complementary fields, such as GPU acceleration and deep neural networks, as well as cross-disciplinary influences from domains like speech recognition. However, during this rapid advancement, certain methods within nanopore sequencing remain relatively unexplored. This oversight has the potential to create bottlenecks in the technology's further development.

In this thesis work, we delve into these uncharted areas, seeking to fill critical gaps and branch the technology into new frontiers. Our objective is to unleash its potential, enabling further breakthroughs in genomic research and beyond.

Through our research, we have developed two novel algorithms and two innovative models tailored to address these under-explored aspects of nanopore sequencing. The two algorithms, the GMBS and the LFBS, both belonging to the MBS group, offer innovative solutions to the challenging decoding problems inherent in HHMMs. They are two distinct variations tailored to different scenarios. While the GMBS is specifically suited for decoding lengthy sequences, such as those encountered in long-read basecalling, the LFBS is optimized for parallel programming and excels in processing short-length sequences. 

The two innovative models developed in this research, each leveraging variations of HHMMs and employing an end-to-end approach, exhibit distinguished structures. The first model, a hybrid of EDHMM and DNN, showcases the effectiveness of integrating both knowledge-driven and data-driven techniques. In contrast, the second model, a custom-designed Helicase HMM, draws inspiration from pioneering studies on motor proteins found in sequencing devices. With its elaborate hierarchical state architecture boasting over five million emission states, this model offers a comprehensive feature space comparable to its predecessor. 

Abstract [sv]

Inom mindre än fyra decennier har nanopore-sekvenseringsteknologin accelererat från en otrolig idé som skissades i en antckningsbok till en avgörande teknologi som bidragit till den kompletta sekvensensieringen av det mänskliga genomet. Denna snabba utveckling, särskilt under de senaste åren, drivs inte bara av innovation kring nanoporer utan också av synergistiska framsteg inom kompletterande områden, såsom GPU-acceleration och djupa neurala nätverk, samt tvärvetenskapligt inflytande från domäner som taligenkänning. Under denna snabba utveckling har do vissa metoder inom nanopore sekvensering förblivit relativt outforskade. Detta förbiseende riskerar att skapa flaskhalsar i teknologins vidareutveckling.

I denna avhandling utforskar vi dessa outforskade områden, i syfte att fylla kritiska luckor och utveckla tekniken mot nya fronter. Vårt mål är att frigöra dess fulla potential och möjliggöra ytterligare genombrott inom genomisk forskning och därutöver.

Som del av vår forskning har vi utvecklat två nya algoritmer och två innovativa modeller anpassade för att adressera dessa underutredda aspekter av nanopore sekvensering. De två algoritmerna, GMBS och LFBS, som är instanser av det mer generella ramverket av MBS-algorithmer (\emph{eng.} marginalised beam search), erbjuder innovativa lösningar på de utmanande avkodningsproblemen som är inneboende i HHMM:er. De är två distinkta variationer anpassade för olika scenarier. Medan GMBS är speciellt lämpad för avkodning av långa sekvenser, såsom de som stöts på vid läsning av långa sekvenser, är LFBS optimerad för parallell programmering och utmärker sig i bearbetning av korta sekvenser.

De två innovativa modellerna som utvecklats i denna forskning, vilka båda utnyttjar variationer av HHMM:er och använder en ``end-to-end''-ansats, uppvisar distinkta strukturer. Den första modellen, en hybrid av EDHMM och DNN, visar effektiviteten av att integrera både kunskapsdrivna och datadrivna tekniker. I kontrast till detta, drar den andra modellen, en anpassad Helicase HMM, inspiration från pionjärstudier om motorproteiner som finns i sekvenseringsenheter. Med sin detaljerade hierarkiska tillståndsarkitektur med över fem miljoner emissivtillstånd, erbjuder denna modell en omfattande egenskapsrymd jämfört med sina föregångare.

Place, publisher, year, edition, pages
KTH Royal Institute of Technology, 2024. , p. 147
Series
TRITA-EECS-AVL ; 2024:40
National Category
Bioinformatics (Computational Biology)
Research subject
Electrical Engineering
Identifiers
URN: urn:nbn:se:kth:diva-346051ISBN: 978-91-8040-914-8 (print)OAI: oai:DiVA.org:kth-346051DiVA, id: diva2:1855511
Public defence
2024-05-24, F3, Lindstedtsvägen 26, Stockholm, 13:00 (English)
Opponent
Supervisors
Note

QC 20240502

Hybrid attendance: https://kth-se.zoom.us/j/66248893067

Available from: 2024-05-02 Created: 2024-05-01 Last updated: 2024-06-05Bibliographically approved
List of papers
1. Marginalized Beam Search Algorithms for Hierarchical HMMs
Open this publication in new window or tab >>Marginalized Beam Search Algorithms for Hierarchical HMMs
2023 (English)Other (Other academic)
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:kth:diva-346020 (URN)
Available from: 2024-04-29 Created: 2024-04-29 Last updated: 2024-05-01Bibliographically approved
2. Lokatt: a hybrid DNA nanopore basecaller with an explicit duration hidden Markov model and a residual LSTM network
Open this publication in new window or tab >>Lokatt: a hybrid DNA nanopore basecaller with an explicit duration hidden Markov model and a residual LSTM network
2023 (English)In: BMC Bioinformatics, E-ISSN 1471-2105, Vol. 24, no 1, article id 461Article in journal (Refereed) Published
Abstract [en]

BackgroundBasecalling long DNA sequences is a crucial step in nanopore-based DNA sequencing protocols. In recent years, the CTC-RNN model has become the leading basecalling model, supplanting preceding hidden Markov models (HMMs) that relied on pre-segmenting ion current measurements. However, the CTC-RNN model operates independently of prior biological and physical insights.ResultsWe present a novel basecaller named Lokatt: explicit duration Markov model and residual-LSTM network. It leverages an explicit duration HMM (EDHMM) designed to model the nanopore sequencing processes. Trained on a newly generated library with methylation-free Ecoli samples and MinION R9.4.1 chemistry, the Lokatt basecaller achieves basecalling performances with a median single read identity score of 0.930, a genome coverage ratio of 99.750%, on par with existing state-of-the-art structure when trained on the same datasets.ConclusionOur research underlines the potential of incorporating prior knowledge into the basecalling processes, particularly through integrating HMMs and recurrent neural networks. The Lokatt basecaller showcases the efficacy of a hybrid approach, emphasizing its capacity to achieve high-quality basecalling performance while accommodating the nuances of nanopore sequencing. These outcomes pave the way for advanced basecalling methodologies, with potential implications for enhancing the accuracy and efficiency of nanopore-based DNA sequencing protocols.

Place, publisher, year, edition, pages
Springer Nature, 2023
Keywords
Basecalling, HMM, LSTM, Nanopore sequencing
National Category
Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:kth:diva-341527 (URN)10.1186/s12859-023-05580-x (DOI)001115621100003 ()38062356 (PubMedID)2-s2.0-85178887529 (Scopus ID)
Note

QC 20231222

Available from: 2023-12-22 Created: 2023-12-22 Last updated: 2024-05-01Bibliographically approved
3. Modelling the nanopore sequencing process with Helicase HMMs
Open this publication in new window or tab >>Modelling the nanopore sequencing process with Helicase HMMs
(English)Manuscript (preprint) (Other academic)
National Category
Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:kth:diva-346049 (URN)
Note

QC 20240508

Available from: 2024-05-01 Created: 2024-05-01 Last updated: 2024-05-08Bibliographically approved

Open Access in DiVA

thesis(41162 kB)129 downloads
File information
File name INSIDE01.pdfFile size 41162 kBChecksum SHA-512
56801dc4d4d52c5a1ef6cf0447c094e67f6307ebf7a4c57512dab24cef839e0957a229f39beb7f1555afb4a916efcd8455ac782c70bde4ec29f365cf4772b105
Type insideMimetype application/pdf

Search in DiVA

By author/editor
Xu, Xuechun
By organisation
Information Science and Engineering
Bioinformatics (Computational Biology)

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

isbn
urn-nbn

Altmetric score

isbn
urn-nbn
Total: 1514 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf