kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Publications (10 of 41) Show all publications
Westrin, K. J., Kretzschmar, W. W. & Emanuelsson, O. (2024). ClusTrast: a short read de novo transcript isoform assembler guided by clustered contigs. BMC Bioinformatics, 25(1), Article ID 54.
Open this publication in new window or tab >>ClusTrast: a short read de novo transcript isoform assembler guided by clustered contigs
2024 (English)In: BMC Bioinformatics, E-ISSN 1471-2105, Vol. 25, no 1, article id 54Article in journal (Refereed) Published
Abstract [en]

Background: Transcriptome assembly from RNA-sequencing data in species without a reliable reference genome has to be performed de novo, but studies have shown that de novo methods often have inadequate ability to reconstruct transcript isoforms. We address this issue by constructing an assembly pipeline whose main purpose is to produce a comprehensive set of transcript isoforms. Results: We present the de novo transcript isoform assembler ClusTrast, which takes short read RNA-seq data as input, assembles a primary assembly, clusters a set of guiding contigs, aligns the short reads to the guiding contigs, assembles each clustered set of short reads individually, and merges the primary and clusterwise assemblies into the final assembly. We tested ClusTrast on real datasets from six eukaryotic species, and showed that ClusTrast reconstructed more expressed known isoforms than any of the other tested de novo assemblers, at a moderate reduction in precision. For recall, ClusTrast was on top in the lower end of expression levels (<15% percentile) for all tested datasets, and over the entire range for almost all datasets. Reference transcripts were often (35–69% for the six datasets) reconstructed to at least 95% of their length by ClusTrast, and more than half of reference transcripts (58–81%) were reconstructed with contigs that exhibited polymorphism, measuring on a subset of reliably predicted contigs. ClusTrast recall increased when using a union of assembled transcripts from more than one assembly tool as primary assembly. Conclusion: We suggest that ClusTrast can be a useful tool for studying isoforms in species without a reliable reference genome, in particular when the goal is to produce a comprehensive transcriptome set with polymorphic variants.

Place, publisher, year, edition, pages
Springer Nature, 2024
Keywords
De novo transcriptome assembly, Guiding contigs, Isoform assembly, Recall/sensitivity, RNA-seq
National Category
Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:kth:diva-343497 (URN)10.1186/s12859-024-05663-3 (DOI)001155913600002 ()38302873 (PubMedID)2-s2.0-85183701397 (Scopus ID)
Note

Not duplicate with DiVA 1810720

QC 20240215

Available from: 2024-02-15 Created: 2024-02-15 Last updated: 2024-02-27Bibliographically approved
Correia, J. C., Jannig, P. R., Gosztyla, M. L., Cervenka, I., Ducommun, S., Praestholm, S. M., . . . Ruas, J. L. (2024). Zfp697 is an RNA- binding protein that regulates skeletal muscle inflammation and remodeling. Proceedings of the National Academy of Sciences of the United States of America, 121(34), Article ID e2319724121.
Open this publication in new window or tab >>Zfp697 is an RNA- binding protein that regulates skeletal muscle inflammation and remodeling
Show others...
2024 (English)In: Proceedings of the National Academy of Sciences of the United States of America, ISSN 0027-8424, E-ISSN 1091-6490, Vol. 121, no 34, article id e2319724121Article in journal (Refereed) Published
Abstract [en]

Skeletal muscle atrophy is a morbidity and mortality risk factor that happens with disuse, chronic disease, and aging. The tissue remodeling that happens during recovery from atrophy or injury involves changes in different cell types such as muscle fibers, and satellite and immune cells. Here, we show that the previously uncharacterized gene and protein Zfp697 is a damage- induced regulator of muscle remodeling. Zfp697/ ZNF697 expression is transiently elevated during recovery from muscle atrophy or injury in mice and humans. Sustained Zfp697 expression in mouse muscle leads to a gene expression signature of chemokine secretion, immune cell recruitment, and extra- cellular matrix remodeling. Notably, although Zfp697 is expressed in several cell types in skeletal muscle, myofiber- specific Zfp697 genetic ablation in mice is sufficient to hinder the inflammatory and regenerative response to muscle injury, compromising functional recovery. We show that Zfp697 is an essential mediator of the interferon gamma response in muscle cells and that it functions primarily as an RNA- interacting protein, with a very high number of miRNA targets. This work identifies Zfp697 as an integrator of cell-cell communication necessary for tissue remodeling and regeneration.

Place, publisher, year, edition, pages
Proceedings of the National Academy of Sciences, 2024
Keywords
skeletal muscle, Zfp697, muscle atrophy, inflammation, RNA- binding protein
National Category
Medical and Health Sciences
Identifiers
urn:nbn:se:kth:diva-357794 (URN)10.1073/pnas.2319724121 (DOI)001369092900002 ()39141348 (PubMedID)2-s2.0-85201774514 (Scopus ID)
Note

QC 20250120

Available from: 2025-01-07 Created: 2025-01-07 Last updated: 2025-01-20Bibliographically approved
Akhter, S., Westrin, K. J., Zivi, N., Nordal, V., Kretzschmar, W. W., Delhomme, N., . . . Sundström, J. F. (2022). Cone-setting in spruce is regulated by conserved elements of the age-dependent flowering pathway. New Phytologist, 236(5), 1951-1963
Open this publication in new window or tab >>Cone-setting in spruce is regulated by conserved elements of the age-dependent flowering pathway
Show others...
2022 (English)In: New Phytologist, ISSN 0028-646X, E-ISSN 1469-8137, Vol. 236, no 5, p. 1951-1963Article in journal (Refereed) Published
Abstract [en]

Reproductive phase change is well characterized in angiosperm model species, but less studied in gymnosperms. We utilize the early cone-setting acrocona mutant to study reproductive phase change in the conifer Picea abies (Norway spruce), a gymnosperm. The acrocona mutant frequently initiates cone-like structures, called transition shoots, in positions where wild-type P. abies always produces vegetative shoots. We collect acrocona and wild-type samples, and RNA-sequence their messenger RNA (mRNA) and microRNA (miRNA) fractions. We establish gene expression patterns and then use allele-specific transcript assembly to identify mutations in acrocona. We genotype a segregating population of inbred acrocona trees. A member of the SQUAMOSA BINDING PROTEIN-LIKE (SPL) gene family, PaSPL1, is active in reproductive meristems, whereas two putative negative regulators of PaSPL1, miRNA156 and the conifer specific miRNA529, are upregulated in vegetative and transition shoot meristems. We identify a mutation in a putative miRNA156/529 binding site of the acrocona PaSPL1 allele and show that the mutation renders the acrocona allele tolerant to these miRNAs. We show co-segregation between the early cone-setting phenotype and trees homozygous for the acrocona mutation. In conclusion, we demonstrate evolutionary conservation of the age-dependent flowering pathway and involvement of this pathway in regulating reproductive phase change in the conifer P. abies. 

Place, publisher, year, edition, pages
Wiley, 2022
Keywords
cone-setting, flowering, gymnosperm, Picea abies, reproductive development, SPL-gene family, transcriptome, allele, angiosperm, coniferous tree, gene expression, genotype, mutation, phenotype, RNA
National Category
Evolutionary Biology
Identifiers
urn:nbn:se:kth:diva-327269 (URN)10.1111/nph.18449 (DOI)000854207600001 ()36076311 (PubMedID)2-s2.0-85138178292 (Scopus ID)
Note

QC 20230523

Available from: 2023-05-23 Created: 2023-05-23 Last updated: 2023-11-09Bibliographically approved
Armenteros, J. J., Salvatore, M., Emanuelsson, O., Winther, O., von Heijne, G., Elofsson, A. & Nielsen, H. (2019). Detecting sequence signals in targeting peptides using deep learning. Life Science Alliance, 2(5), Article ID UNSP e201900429.
Open this publication in new window or tab >>Detecting sequence signals in targeting peptides using deep learning
Show others...
2019 (English)In: Life Science Alliance, E-ISSN 2575-1077, Vol. 2, no 5, article id UNSP e201900429Article in journal (Refereed) Published
Abstract [en]

In bioinformatics, machine learning methods have been used to predict features embedded in the sequences. In contrast to what is generally assumed, machine learning approaches can also provide new insights into the underlying biology. Here, we demonstrate this by presenting TargetP 2.0, a novel state-of-the-art method to identify N-terminal sorting signals, which direct proteins to the secretory pathway, mitochondria, and chloroplasts or other plastids. By examining the strongest signals from the attention layer in the network, we find that the second residue in the protein, that is, the one following the initial methionine, has a strong influence on the classification. We observe that two-thirds of chloroplast and thylakoid transit peptides have an alanine in position 2, compared with 20% in other plant proteins. We also note that in fungi and single-celled eukaryotes, less than 30% of the targeting peptides have an amino acid that allows the removal of the N-terminal methionine compared with 60% for the proteins without targeting peptide. The importance of this feature for predictions has not been highlighted before.

Place, publisher, year, edition, pages
LIFE SCIENCE ALLIANCE LLC, 2019
National Category
Biological Sciences
Identifiers
urn:nbn:se:kth:diva-264335 (URN)10.26508/lsa.201900429 (DOI)000494674100006 ()31570514 (PubMedID)2-s2.0-85072779066 (Scopus ID)
Note

QC 20191126

Available from: 2019-11-26 Created: 2019-11-26 Last updated: 2022-06-26Bibliographically approved
Akhter, S., Kretzschmar, W. W., Nordal, V., Delhomme, N., Street, N. R., Nilsson, O., . . . Sundström, J. F. (2018). Integrative Analysis of Three RNA Sequencing Methods Identifies Mutually Exclusive Exons of MADS-Box Isoforms During Early Bud Development in Picea abies. Frontiers in Plant Science, 9, Article ID 1625.
Open this publication in new window or tab >>Integrative Analysis of Three RNA Sequencing Methods Identifies Mutually Exclusive Exons of MADS-Box Isoforms During Early Bud Development in Picea abies
Show others...
2018 (English)In: Frontiers in Plant Science, E-ISSN 1664-462X, Vol. 9, article id 1625Article in journal (Refereed) Published
Abstract [en]

Recent efforts to sequence the genomes and transcriptomes of several gymnosperm species have revealed an increased complexity in certain gene families in gymnosperms as compared to angiosperms. One example of this is the gymnosperm sister Glade to angiosperm TM3-like MADS-box genes, which at least in the conifer lineage has expanded in number of genes. We have previously identified a member of this subclade, the conifer gene DEFICIENS AGAMOUS LIKE 19 (DAL19), as being specifically upregulated in cone-setting shoots. Here, we show through Sanger sequencing of mRNA-derived cDNA and mapping to assembled conifer genomic sequences that DAL19 produces six mature mRNA splice variants in Picea abies. These splice variants use alternate first and last exons, while their four central exons constitute a core region present in all six transcripts. Thus, they are likely to be transcript isoforms. Quantitative Real-Time PCR revealed that two mutually exclusive first DAL19 exons are differentially expressed across meristems that will form either male or female cones, or vegetative shoots. Furthermore, mRNA in situ hybridization revealed that two mutually exclusive last DAL19 exons were expressed in a cell-specific pattern within bud meristems. Based on these findings in DAL19, we developed a sensitive approach to transcript isoform assembly from short-read sequencing of mRNA. We applied this method to 42 putative MADS-box core regions in P abies, from which we assembled 1084 putative transcripts. We manually curated these transcripts to arrive at 933 assembled transcript isoforms of 38 putative MADS-box genes. 152 of these isoforms, which we assign to 28 putative MADS-box genes, were differentially expressed across eight female, male, and vegetative buds. We further provide evidence of the expression of 16 out of the 38 putative MADS-box genes by mapping PacBio Iso-Seq circular consensus reads derived from pooled sample sequencing to assembled transcripts. In summary, our analyses reveal the use of mutually exclusive exons of MADS-box gene isoforms during early bud development in P. abies, and we find that the large number of identified MADS-box transcripts in P. abies results not only from expansion of the gene family through gene duplication events but also from the generation of numerous splice variants.

Place, publisher, year, edition, pages
Frontiers Media S.A., 2018
Keywords
Picea abies, MADS-box genes, cone development, De Bruijn assembly, transcript isoforms, RNA sequencing, DAL19
National Category
Genetics and Genomics
Identifiers
urn:nbn:se:kth:diva-239473 (URN)10.3389/fpls.2018.01625 (DOI)000449948700001 ()30483285 (PubMedID)2-s2.0-85058812295 (Scopus ID)
Funder
Knut and Alice Wallenberg FoundationVinnova
Note

QC 20181126

Available from: 2018-11-26 Created: 2018-11-26 Last updated: 2025-02-07Bibliographically approved
Reimegård, J., Kundu, S., Pendle, A., Irish, V. F., Shaw, P., Nakayama, N., . . . Emanuelsson, O. (2017). Genome-wide identification of physically clustered genes suggests chromatin-level co-regulation in male reproductive development in Arabidopsis thaliana. Nucleic Acids Research
Open this publication in new window or tab >>Genome-wide identification of physically clustered genes suggests chromatin-level co-regulation in male reproductive development in Arabidopsis thaliana
Show others...
2017 (English)In: Nucleic Acids Research, ISSN 0305-1048, E-ISSN 1362-4962Article in journal (Refereed) Published
Abstract [en]

Co-expression of physically linked genes occurs surprisingly frequently in eukaryotes. Such chromosomal clustering may confer a selective advantage as it enables coordinated gene regulation at the chromatin level. We studied the chromosomal organization of genes involved in male reproductive development in Arabidopsis thaliana. We developed an in-silico tool to identify physical clusters of co-regulated genes from gene expression data. We identified 17 clusters (96 genes) involved in stamen development and acting downstream of the transcriptional activator MS1 (MALE STERILITY 1), which contains a PHD domain associated with chromatin re-organization. The clusters exhibited little gene homology or promoter element similarity, and largely overlapped with reported repressive histone marks. Experiments on a subset of the clusters suggested a link between expression activation and chromatin conformation: qRT-PCR and mRNA in situ hybridization showed that the clustered genes were up-regulated within 48 h after MS1 induction; out of 14 chromatin-remodeling mutants studied, expression of clustered genes was consistently down-regulated only in hta9/hta11, previously associated with metabolic cluster activation; DNA fluorescence in situ hybridization confirmed that transcriptional activation of the clustered genes was correlated with open chromatin conformation. Stamen development thus appears to involve transcriptional activation of physically clustered genes through chromatin de-condensation.

Place, publisher, year, edition, pages
Oxford University Press, 2017
National Category
Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:kth:diva-203892 (URN)10.1093/nar/gkx087 (DOI)000398376200035 ()28175342 (PubMedID)2-s2.0-85019930165 (Scopus ID)
Note

QC 20170320

Available from: 2017-03-20 Created: 2017-03-20 Last updated: 2024-03-18Bibliographically approved
Edsgärd, D., Iglesias, M. J., Reilly, S.-J., Hamsten, A., Tornvall, P., Odeberg, J. & Emanuelsson, O. (2016). GeneiASE: Detection of condition-dependent and static allele-specific expression from RNA-seq data without haplotype information. Scientific Reports, 6, Article ID 21134.
Open this publication in new window or tab >>GeneiASE: Detection of condition-dependent and static allele-specific expression from RNA-seq data without haplotype information
Show others...
2016 (English)In: Scientific Reports, E-ISSN 2045-2322, Vol. 6, article id 21134Article in journal (Refereed) Published
Abstract [en]

Allele-specific expression (ASE) is the imbalance in transcription between maternal and paternal alleles at a locus and can be probed in single individuals using massively parallel DNA sequencing technology. Assessing ASE within a single sample provides a static picture of the ASE, but the magnitude of ASE for a given transcript may vary between different biological conditions in an individual. Such condition-dependent ASE could indicate a genetic variation with a functional role in the phenotypic difference. We investigated ASE through RNA-sequencing of primary white blood cells from eight human individuals before and after the controlled induction of an inflammatory response, and detected condition-dependent and static ASE at 211 and 13021 variants, respectively. We developed a method, GeneiASE, to detect genes exhibiting static or condition-dependent ASE in single individuals. GeneiASE performed consistently over a range of read depths and ASE effect sizes, and did not require phasing of variants to estimate haplotypes. We observed condition-dependent ASE related to the inflammatory response in 19 genes, and static ASE in 1389 genes. Allele-specific expression was confirmed by validation of variants through real-time quantitative RT-PCR, with RNA-seq and RT-PCR ASE effect-size correlations r = 0.67 and r = 0.94 for static and condition-dependent ASE, respectively.

Place, publisher, year, edition, pages
Nature Publishing Group, 2016
National Category
Engineering and Technology
Identifiers
urn:nbn:se:kth:diva-183310 (URN)10.1038/srep21134 (DOI)000370363200001 ()26887787 (PubMedID)2-s2.0-84958999591 (Scopus ID)
Funder
Science for Life Laboratory - a national resource center for high-throughput molecular bioscience
Note

QC 20160309

Available from: 2016-03-09 Created: 2016-03-07 Last updated: 2024-03-15Bibliographically approved
Sigurgeirsson, B., Emanuelsson, O. & Lundeberg, J. (2014). Analysis of stranded information using an automated procedure for strand specific RNA sequencing. BMC Genomics, 15(1), Article ID 631.
Open this publication in new window or tab >>Analysis of stranded information using an automated procedure for strand specific RNA sequencing
2014 (English)In: BMC Genomics, E-ISSN 1471-2164, Vol. 15, no 1, article id 631Article in journal (Refereed) Published
Abstract [en]

Background: Strand specific RNA sequencing is rapidly replacing conventional cDNA sequencing as an approach for assessing information about the transcriptome. Alongside improved laboratory protocols the development of bioinformatical tools is steadily progressing. In the current procedure the Illumina TruSeq library preparation kit is used, along with additional reagents, to make stranded libraries in an automated fashion which are then sequenced on Illumina HiSeq 2000. By the use of freely available bioinformatical tools we show, through quality metrics, that the protocol is robust and reproducible. We further highlight the practicality of strand specific libraries by comparing expression of strand specific libraries to non-stranded libraries, by looking at known antisense transcription of pseudogenes and by identifying novel transcription. Furthermore, two ribosomal depletion kits, RiboMinus and RiboZero, are compared and two sequence aligners, Tophat2 and STAR, are also compared. Results: The, non-stranded, Illumina TruSeq kit can be adapted to generate strand specific libraries and can be used to access detailed information on the transcriptome. The RiboZero kit is very effective in removing ribosomal RNA from total RNA and the STAR aligner produces high mapping yield in a short time. Strand specific data gives more detailed and correct results than does non-stranded data as we show when estimating expression values and in assembling transcripts. Even well annotated genomes need improvements and corrections which can be achieved using strand specific data. Conclusions: Researchers in the field should strive to use strand specific data; it allows for more confidence in the data analysis and is less likely to lead to false conclusions. If faced with analysing non-stranded data, researchers should be well aware of the caveats of that approach.

Keywords
Antisense RNA, Bioinformatics, Ribosomal depletion, RNA sequencing, Strand specificity
National Category
Biological Sciences
Identifiers
urn:nbn:se:kth:diva-161768 (URN)10.1186/1471-2164-15-631 (DOI)000209596800001 ()25070246 (PubMedID)2-s2.0-84904776692 (Scopus ID)
Funder
Science for Life Laboratory - a national resource center for high-throughput molecular bioscienceSwedish Research Council
Note

QC 20150317

Available from: 2015-03-17 Created: 2015-03-17 Last updated: 2024-03-18Bibliographically approved
Emanuelsson, O., Arvestad, L. & Käll, L. (2014). Engagera och aktivera studenter med inspiration från konferenser: examination genom poster-presentation. In: Roy Andersson (Ed.), Proceedings 2014, 8:e Pedagogiska inspirationskonferensen 17 december 2014: . Paper presented at LTHs 8:e Pedagogiska Inspirationskonferens, 17 december 2014. Lund
Open this publication in new window or tab >>Engagera och aktivera studenter med inspiration från konferenser: examination genom poster-presentation
2014 (Swedish)In: Proceedings 2014, 8:e Pedagogiska inspirationskonferensen 17 december 2014 / [ed] Roy Andersson, Lund, 2014Conference paper, Published paper (Refereed)
Abstract [sv]

I en forskningsnära kurs om 7.5 hp på master-nivå inom bioinformatikämnet vid KTH består drygt halva kursen av ett projekt som genomförs i grupper om tre studenter. Varje projekt har en egen projektuppgift med inget eller marginellt överlapp med andra gruppers uppgifter. Projekten är så gott som uteslutande baserade på aktuella frågeställningar i lärarteamets egna forskningsgrupper eller deras närhet. Projektet redovisas dels genom en posterpresentation, dels med individuell webbaserad projektdagbok. Vid posterredovisningen, som omfattar tre timmar i slutet av tentamensperioden, är alla kursdeltagare med. Vi försöker i möjligaste mån efterlikna situationen där ett autentiskt forskningsresultat presenteras på en riktig konferens. Varje deltagare (student) förväntas alltså ta del av varje annan grupps poster, på samma sätt som sker vid de flesta vetenskapliga konferenser. Vi genomför en enklare kamratbedömning på posternivå, där varje student ska avge en kort och konfidentiell kommentar om var och en av övriga postrar. Kursens lärare bedömer förstås också postrarna. En av svårigheterna är att sätta individuella betyg. Här använder vi oss av individuella projektdagböcker, som ger vägledning till de olika individernas insatser inom projektet. Vi har provat detta under fyra kursomgångar med som mest sju projekt. Examinationsformen är rolig och motiverande både för studenterna och lärarna.

Place, publisher, year, edition, pages
Lund: , 2014
Keywords
examination, konstruktiv länkning, bioinformatik, projektkurs, poster-presentation
National Category
Other Engineering and Technologies
Research subject
Education and Communication in the Technological Sciences
Identifiers
urn:nbn:se:kth:diva-163041 (URN)
Conference
LTHs 8:e Pedagogiska Inspirationskonferens, 17 december 2014
Projects
Pedagogiska utvecklare vid KTH
Note

QC 20150327

Available from: 2015-03-26 Created: 2015-03-26 Last updated: 2025-02-10Bibliographically approved
Song, Y., Giske, C. G., Gille-Johnson, P., Emanuelsson, O., Lundeberg, J. & Gyarmati, P. (2014). Nuclease-Assisted Suppression of Human DNA Background in Sepsis. PLOS ONE, 9(7), e103610
Open this publication in new window or tab >>Nuclease-Assisted Suppression of Human DNA Background in Sepsis
Show others...
2014 (English)In: PLOS ONE, E-ISSN 1932-6203, Vol. 9, no 7, p. e103610-Article in journal (Refereed) Published
Abstract [en]

Sepsis is a severe medical condition characterized by a systemic inflammatory response of the body caused by pathogenic microorganisms in the bloodstream. Blood or plasma is typically used for diagnosis, both containing large amount of human DNA, greatly exceeding the DNA of microbial origin. In order to enrich bacterial DNA, we applied the C(0)t effect to reduce human DNA background: a model system was set up with human and Escherichia coli (E. coli) DNA to mimic the conditions of bloodstream infections; and this system was adapted to plasma and blood samples from septic patients. As a consequence of the C(0)t effect, abundant DNA hybridizes faster than rare DNA. Following denaturation and re-hybridization, the amount of abundant DNA can be decreased with the application of double strand specific nucleases, leaving the non-hybridized rare DNA intact. Our experiments show that human DNA concentration can be reduced approximately 100,000-fold without affecting the E. coli DNA concentration in a model system with similarly sized amplicons. With clinical samples, the human DNA background was decreased 100-fold, as bacterial genomes are approximately 1,000-fold smaller compared to the human genome. According to our results, background suppression can be a valuable tool to enrich rare DNA in clinical samples where a high amount of background DNA can be found.

National Category
Analytical Chemistry
Identifiers
urn:nbn:se:kth:diva-143193 (URN)10.1371/journal.pone.0103610 (DOI)000340028800068 ()25076135 (PubMedID)2-s2.0-84905054440 (Scopus ID)
Note

Updated from manuscript to article in journal.

QC 20140912

Available from: 2014-03-18 Created: 2014-03-18 Last updated: 2024-03-18Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-8879-9245

Search in DiVA

Show all publications