kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
The art of transcriptome reconstruction: with applications in Picea abies (L.) H. Karst
KTH, School of Engineering Sciences in Chemistry, Biotechnology and Health (CBH), Gene Technology. (Emanuelsson Lab)ORCID iD: 0000-0002-6937-9245
2023 (English)Doctoral thesis, comprehensive summary (Other academic)Alternative title
Konsten att rekonstruera transkriptom : med tillämpningar i rödgran (Swedish)
Abstract [en]

Transcriptome reconstruction is an important component in the bioinformatical part of transcriptome studies. When a reference genome is missing, highly fragmented or incomplete, a de novo transcriptome assembly is the transcriptome reconstruction approach of choice, since in such situations, a simple alignment (or mapping) would not necessarily give all theinformation concerning splice junctions, isoforms or even the full extent of the gene. Several methods for de novo transcriptome assembly have been suggested, but many of these methods lack sufficient ability to recover isoforms or are memory intense, which requires themethods to be executed on computing clusters.

One species, whose published reference genome is highly fragmented, is the Norway spruce (Picea abies (L.) H. Karst.) – a conifer, very important for Swedish forestry ande conomy, but with a long juvenile phase and irregular cone setting, the demand of cultivated seeds is larger than the supply. Thus, there is a desire to understand the molecular biology behind the cone setting in P. abies, not least regarding gene expression and its regulation. This doctoral dissertation addresses these problems by describing the biological background in general, followed by an introduction to theoretical computational problems relatedto the methods applied for transcriptome reconstruction, which then are described in depth themselves, as is P. abies.

Paper I uses a novel de novo assembler to detect connections between scaffolds in the P. abies genome, and also studies P. abies var acrocona, a mutant with shorter juvenile phase and more regular cone setting than the wild type, in order to detect how cone setting is initiated. By means of allele-specific expression analysis, this study detects a SNP ina miRNA binding site on a novel gene, a mutation which is coherent with the acroconaphenotype.

Paper II and paper III both introduce one novel de novo transcriptome assembler each: Paper II describes the assembly method applied in paper I, with the focus torecover a comprehensive list of isoforms. It accomplished this, thus providing higher recallthan the other tested assemblers, but with an increased use of computational resources. In turn, paper III introduces a lightweight assembly method, which is the first assembly method employing the ant colony system (ACS) meta-heuristic. It is more rapid than theother tested assemblers and requires less memory (never more than half of the second most memory efficient method), but provides low recall.

Paper IV applies reference based transcriptome assembly to improve the gene annotation of a new, chromosome-scale, P. abies reference genome, which is being prepared at the moment. This study pinpoints the locations in the assembled genome of six previously described genes, but which the annotation missed – until now. Furthermore, for two annotatedgenes, this study found and verified one novel transcript isoform each.

Paper V studies the natural variation in cone setting in P. abies, and also the gene expression pre and post treatment of gibberellic acid (GA), known to stimulate floweringin plants. P. abies genotypes with lower cone setting ability turns out to get more genesactivated post GA treatment, compared to genotypes with higher cone setting ability.

Abstract [sv]

Transkriptomrekonstruktion är en viktig beståndsdel i den bioinformatiska avdelningen av transkriptomstudier. När ett referensgenom saknas, är kraftigt fragmenterat eller ofullständigt, så måste sammansättningen av transkriptomet ske de novo, ty i dessa situationer skulle inte en vanlig inpassning (eller mappning) nödvändigtvis ge all information om splitsningsplatser, isoformer eller ens genens fullständiga omfattning.  Flertalet metoder för sammansättning av transkriptom de novo har föreslagits, men många av dessa saknar förmåga att återskapa isoformer eller är minnesintensiva, vilket kräver att metoderna exekveras på datorkluster.

En art, vars referensgenom är kraftigt fragmenterat, är rödgran (Picea abies (L.) H. Karst.) - ett barrträd som är mycket viktigt för svenskt skogsbruk och svensk ekonomi, men med en lång uppväxtsfas och oregelbunden kottsättning så är efterfrågan på förädlade fröer större än utbudet. Således finns ett intresse att förstå kottsättningens bakomliggande molekylärbiologi, inte minst med avseenede på vilka gener som uttrycks och hur de regleras.

Denna doktorsavhandling behandlar dessa problem medelst beskrivande av den allmänna biologiska bakgrunden, följt av en introduktion av problem från teoretisk datalogi, vilka är relaterade till metoder för sammansättning av transkriptom, som i sin tur själva är beskrivna i detalj därefter, liksom också rödgran är.

Artikel I tillämpar en ny sammasättningsmetod för att upptäcka kopplingar mellan olika fragment i grangenomet, men studerar även en mutant: P. abies var acrocona (kottegran), vilken har kortare uppväxtsfas och mer regelbunden kottsättning än vildtypen, för att avgöra hur kottsättning initieras. Med hjälp av allelspecifik uttrycksanalys hittar denna stuide en SNP i ett bindningsställe för miRNA i en tidigare ostuderad gen, vilken visar sig sammanhängande med kottegranens fenotyp.

Artikel II och artikel III presenterar varsin ny metod för de novo sammansättning av transkriptom: Artikel II beskriver den sammansättningsmetod som tillämpas i artikel I, med fokus på att återskapa en omfattande lista av isoformer. Detta erhålls, således uppnår metoden en högre grad av sensitivitet än de andra testade metoderna, men till en kostnad av ett ökat behov av beräkningsresurser. I sin tur presenterar artikel III en lättviktig sammansättningsmetod, den första i sitt slag som tillämpar metaheuristiken myrkolonisystem (ant colony system). Metoden är snabbare och minnessnålare än andra testade metoder (använder aldrig mer än hälften av det minne som krävs av den näst minnessnålaste), men ger låg sensitivitet.

Artikel IV tillämpar referensbaserad sammansättning av transkriptom för att förbättra annoteringen av ett nytt, kromosomskaligt referensgenom för rödgran, vilket förberedes i skrivande stund. Denna studie lokaliserar sex tidigare beskrivna gener i det nya genomet, men som missats av annoteringen - tills nu. Vidare får två annoterade gener varsin ny isoform, som hittats och verifierats i denna studie.

Artikel V studerar naturlig variation av kottsättning i rödgran, samt genuttrycksnivåer före och efter behandling av gibberellinsyra (GA), känd att stimulera blomning. Genotyper av rödgran med lägre kottsättningsförmåga tycks få fler gener aktiverade efter GA-behandling, jämfört med genotyper med högre kottsättningsförmåga.

Place, publisher, year, edition, pages
Stockholm: Kungliga Tekniska högskolan, 2023. , p. iv, 55
Series
TRITA-CBH-FOU ; 2023:55
Keywords [en]
transcriptome reconstruction, transcriptome assembly, transcript isoform detection, Picea abies, acrocona, cone setting, differential expression, gene families, miRNA, quality assessment
Keywords [sv]
transkriptomrekonstruktion, sammansättning av transkriptom, detektering av transkriptisoform, rödgran, kottegran, kottsättning, differentiellt uttryck, genfamiljer, kvalitetsutvärdering, miRNA
National Category
Bioinformatics (Computational Biology) Bioinformatics and Computational Biology
Research subject
Biotechnology
Identifiers
URN: urn:nbn:se:kth:diva-339409ISBN: 978-91-8040-774-8 (print)OAI: oai:DiVA.org:kth-339409DiVA, id: diva2:1810977
Public defence
2023-12-08, Air & Fire, Science For Life Laboratory, Tomtebodavägen 23a, https://kth-se.zoom.us/j/69664293949, Solna, 10:00 (English)
Opponent
Supervisors
Note

QC 20231110

Available from: 2023-11-10 Created: 2023-11-09 Last updated: 2025-10-30Bibliographically approved
List of papers
1. Cone-setting in spruce is regulated by conserved elements of the age-dependent flowering pathway
Open this publication in new window or tab >>Cone-setting in spruce is regulated by conserved elements of the age-dependent flowering pathway
Show others...
2022 (English)In: New Phytologist, ISSN 0028-646X, E-ISSN 1469-8137, Vol. 236, no 5, p. 1951-1963Article in journal (Refereed) Published
Abstract [en]

Reproductive phase change is well characterized in angiosperm model species, but less studied in gymnosperms. We utilize the early cone-setting acrocona mutant to study reproductive phase change in the conifer Picea abies (Norway spruce), a gymnosperm. The acrocona mutant frequently initiates cone-like structures, called transition shoots, in positions where wild-type P. abies always produces vegetative shoots. We collect acrocona and wild-type samples, and RNA-sequence their messenger RNA (mRNA) and microRNA (miRNA) fractions. We establish gene expression patterns and then use allele-specific transcript assembly to identify mutations in acrocona. We genotype a segregating population of inbred acrocona trees. A member of the SQUAMOSA BINDING PROTEIN-LIKE (SPL) gene family, PaSPL1, is active in reproductive meristems, whereas two putative negative regulators of PaSPL1, miRNA156 and the conifer specific miRNA529, are upregulated in vegetative and transition shoot meristems. We identify a mutation in a putative miRNA156/529 binding site of the acrocona PaSPL1 allele and show that the mutation renders the acrocona allele tolerant to these miRNAs. We show co-segregation between the early cone-setting phenotype and trees homozygous for the acrocona mutation. In conclusion, we demonstrate evolutionary conservation of the age-dependent flowering pathway and involvement of this pathway in regulating reproductive phase change in the conifer P. abies. 

Place, publisher, year, edition, pages
Wiley, 2022
Keywords
cone-setting, flowering, gymnosperm, Picea abies, reproductive development, SPL-gene family, transcriptome, allele, angiosperm, coniferous tree, gene expression, genotype, mutation, phenotype, RNA
National Category
Evolutionary Biology
Identifiers
urn:nbn:se:kth:diva-327269 (URN)10.1111/nph.18449 (DOI)000854207600001 ()36076311 (PubMedID)2-s2.0-85138178292 (Scopus ID)
Note

QC 20230523

Available from: 2023-05-23 Created: 2023-05-23 Last updated: 2023-11-09Bibliographically approved
2. ClusTrast: a short read de novo transcript isoform assembler guided by clustered contigs
Open this publication in new window or tab >>ClusTrast: a short read de novo transcript isoform assembler guided by clustered contigs
(English)Manuscript (preprint) (Other academic)
Abstract [en]

Background: Transcriptome assembly from RNA-sequencing data in species without a reliable referencegenome has to be performed de novo, but studies have shown that de novo methods often have inadequateability to reconstruct transcript isoforms. We address this issue by constructing an assembly pipeline whosemain purpose is to produce a comprehensive set of transcript isoforms.

Results: We present the de novo transcript isoform assembler ClusTrast, which takes short read RNA-seq dataas input, assembles a primary assembly, clusters a set of guiding contigs, aligns the short reads to the guidingcontigs, assembles each clustered set of short reads individually, and merges the primary and clusterwiseassemblies into the final assembly. We tested ClusTrast on real datasets from six eukaryotic species, andshowed that ClusTrast reconstructed more expressed known isoforms than any of the other tested de novoassemblers, at a moderate reduction in precision. For recall, ClusTrast was on top in the lower end ofexpression levels (<15% percentile) for all tested datasets, and over the entire range for almost all datasets.Reference transcripts were often (35–69% for the six datasets) reconstructed to at least 95% of their length byClusTrast, and more than half of reference transcripts (58–81%) were reconstructed with contigs that exhibitedpolymorphism, measuring on a subset of reliably predicted contigs. ClusTrast recall increased when using aunion of assembled transcripts from more than one assembly tool as primary assembly.

Conclusion: We suggest that ClusTrast can be a useful tool for studying isoforms in species without a reliablereference genome, in particular when the goal is to produce a comprehensive transcriptome set withpolymorphic variants.

Keywords
de novo transcriptome assembly; guiding contigs; isoform assembly; RNA-seq; recall/sensitivity
National Category
Bioinformatics and Computational Biology Bioinformatics (Computational Biology)
Research subject
Biotechnology
Identifiers
urn:nbn:se:kth:diva-339407 (URN)
Note

QC 20231110

Available from: 2023-11-09 Created: 2023-11-09 Last updated: 2025-02-05Bibliographically approved
3. An experimental application of the meta-heuristic Ant Colony System for low memory intense de novo transcriptome assembly of short reads
Open this publication in new window or tab >>An experimental application of the meta-heuristic Ant Colony System for low memory intense de novo transcriptome assembly of short reads
(English)Manuscript (preprint) (Other academic)
Abstract [en]

The sequence assembly problem, known to be NP-hard, is one of the fundamental problems in bioinformatics, not the least when studying genomes or transcriptomes. Several different methods have been proposed, both with and without reference guidance, but most of these methods are memory intense.

Here, we attempt to prove the concept of applying Ant Colony System (ACS), a swarm intelligence meta-heuristic, to the problem of de novo transcriptome assembly. First, we show how the assembly problem can be reduced to a set of established NP-hard problems, where ACS has previously been applied. Next, we develop an ACS-based de novo transcriptome assembler. Finally, we benchmark the novel assembler with a set of newly developed and/or widely used transcriptome assemblers and find that the ACS approach provides the most memory efficient assembly of the tested methods, using at most 49\% of the memory required by any other method, with a moderate reduction in precision and an appreciable reduction in recall.  This suggests that memory efficient de novo transcriptome assembly using ACS is a working concept,, where further improvements are desired. The ACS approach may potentially also be useful in other assembly areas.

Keywords
Transcriptome assembly, RNA-seq data, ant colony optimization, de Bruijn graph, meta-heuristic, graph theory
National Category
Bioinformatics and Computational Biology Bioinformatics (Computational Biology)
Research subject
Biotechnology; Computer Science
Identifiers
urn:nbn:se:kth:diva-339404 (URN)
Note

QC 20231110

Available from: 2023-11-09 Created: 2023-11-09 Last updated: 2025-02-05Bibliographically approved
4. Improving the annotation of three reproductive gene families in Picea abies
Open this publication in new window or tab >>Improving the annotation of three reproductive gene families in Picea abies
Show others...
(English)Manuscript (preprint) (Other academic)
Abstract [en]

The gene families MADS-box, FT and SPL are all known to be related to cone setting inconifers. Since the genome of most conifers are large and repetitive, no full in-depth studyof these gene families (or any other gene family in conifers) has been possible.

With the advent of a new, near-complete, reference genome assembly of Picea abies, weperformed an in-depth analysis of these gene families, in order to find novel genes withinthese families and simultaneously verify the quality of the newly assembled reference genomeand of its corresponding annotation. We used 30 Illumina and two PacBio RNA-seq datasetsfrom cone setting studies and assembled the corresponding transcriptomes, from which weidentified 35 novel MADS-box genes, 5 novel SPL-genes and 5 novel FT-genes among theannotated genes. We also verified the existence of an exon in the MADS-box gene DAL19,previously only detected by assembly, and we pinpointed the genomic location of five previ-ously identified MADS-box genes missing in the new annotation. This study leaves us withthe same number of MADS-box and SPL-genes as was annotated in a recently publishedreference genome in another conifer, Pinus tabuliformis.

National Category
Bioinformatics and Computational Biology Bioinformatics (Computational Biology)
Research subject
Biotechnology
Identifiers
urn:nbn:se:kth:diva-339405 (URN)
Note

QC 20231110

Available from: 2023-11-09 Created: 2023-11-09 Last updated: 2025-02-05Bibliographically approved
5. Natural variation in cone-setting ability in Norway spruce
Open this publication in new window or tab >>Natural variation in cone-setting ability in Norway spruce
Show others...
(English)Manuscript (preprint) (Other academic)
Abstract [en]

The conifer Norway spruce (Picea abies) has an uneven cone-setting and initiates cones onlyevery 3rd to 5th year. The cone-setting is in large synchronised on the population level,resulting in good or bad cone-years. However, variation in cone-setting frequencies andnumber of cones produced per tree can be observed among different genotypes. Here, we aimto study the genetic basis for the natural variation in cone-setting ability. To this end, we haveset up a field trial consisting of 560 ramets representing eighteen P. abies genotypes and havemonitored the cone-setting in these trees since 2008. Using this material, we study the naturalvariation in regulatory elements of a key cone-setting regulator, Picea abies SQUAMOSABINDING PROTEIN LIKE1 (PaSPL1). We have previously established that the combinedaction of miRNA156 and miRNA529 negatively regulates PaSPL1. Here, we use massivelyparallel sequencing to study the expression of these and other microRNAs in the P. abiesgenotypes. While no direct link can be established between miRNA expression levels and cone-setting frequencies, we report on individual genotype differences that grade the relativeimportance of miR156 and miR529 in regulating PaSPL1 transcript levels. We also clone, andSanger sequence the putative promoter of PaSPL1 from the genotypes in the field trial. Byanalysing variation in SNP data, we identify putative Gibberellic Acid (GA) responsiveelements enriched in the genotype with the highest cone-setting frequency. As a complement,we also sequence the transcriptomes of three genotypes before and after GA-inductions. Thisanalysis shows that relatively more genes are activated in a low cone-setting genotype than ina high cone-setting genotype. We also demonstrate that GA-regulated MYB transcriptionfactors are upregulated upon GA induction, providing a possible link between GA-inducedcone-setting and the regulation of PaSPL1 that merits further studies.

Keywords
cone-setting, flowering, gymnosperm, Picea abies, reproductive development, transcriptome
National Category
Bioinformatics and Computational Biology Genetics and Breeding in Agricultural Sciences
Research subject
Biotechnology
Identifiers
urn:nbn:se:kth:diva-339406 (URN)
Note

QC 20231110

Available from: 2023-11-09 Created: 2023-11-09 Last updated: 2025-02-05Bibliographically approved

Open Access in DiVA

Summary(15600 kB)613 downloads
File information
File name FULLTEXT01.pdfFile size 15600 kBChecksum SHA-512
1af62d3edd76c91d5c5f030fe9451cbfc2c4bd101671a70462bd44463332a0e5ae849652be3a4ce3d17983313f49c1bf207ffcf968ae7e08a5dac1c6cdb12714
Type fulltextMimetype application/pdf

Authority records

Westrin, Karl Johan

Search in DiVA

By author/editor
Westrin, Karl Johan
By organisation
Gene Technology
Bioinformatics (Computational Biology)Bioinformatics and Computational Biology

Search outside of DiVA

GoogleGoogle Scholar
Total: 614 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

isbn
urn-nbn

Altmetric score

isbn
urn-nbn
Total: 1394 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf