kth.sePublications KTH
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Data, Geometry and Homology
KTH, School of Engineering Sciences (SCI), Mathematics (Dept.), Algebra, Combinatorics and Topology.ORCID iD: 0000-0002-1513-5069
2026 (English)Doctoral thesis, comprehensive summary (Other academic)Alternative title
Data, Geometri och Homologi (Swedish)
Abstract [en]

Modern datasets are increasingly complex and heterogeneous. Analyzing such data raises fundamental questions about the mathematical spaces in which data should be represented, how objects in these spaces should be compared, and which computable invariants preserve information relevant to a given task. This thesis studies these questions with topological data analysis as a central framework, using homology as a language for describing the geometry of data.

The first part of the thesis develops distances and invariants for spaces of persistence modules. The starting point is categorical: we regard data as objects living in categories equipped with enough algebraic structure to define a notion of size. In particular, abelian categories provide kernels, cokernels, and exact sequences, allowing distances between objects to be constructed from the failure of morphisms to be isomorphisms. Persistence modules form a central example: they encode how homological features appear and disappear along a filtration, and may be viewed as functors from a partially ordered parameter space to vector spaces. Within this framework, we study distances induced by contours, which provide a flexible way of specifying the geometry of the parameter space. This setting leads to compactness results for families of multidimensional persistence modules. In the one-dimensional setting, where a barcode decomposition is available, we develop algebraic Wasserstein distances based on ℓp norms of contour-dependent bar lifetimes. Using these distances, we define Wasserstein stable ranks, stable and computable invariants whose interpretable parameters can be learned for a given task.

The second part of the thesis moves from the mathematical framework to applications in neuroscience, where cellular morphologies provide natural examples of structured geometric data. Microglia and other branched cells can be represented as rooted trees embedded in three-dimensional space, and their morphology can be characterized using topological morphology descriptors. In the morphOMICs pipeline, such descriptors are combined with vectorizations, bootstrapping, dimensionality reduction, and classification in order to map microglial morphology across brain regions and sexes, and through development, disease progression, and experimental perturbations. This gives a data-driven atlas of microglial morphology that avoids relying on preselected scalar morphometric features.

We further introduce the chromatic topological morphology descriptor (chromatic TMD) to study intracellular organization in branched cells. Here a microglial cell is represented by a rooted tree, while CD68-positive and mitochondria organelles are represented by subgraphs of that tree. The inclusion of the organelle subgraph into the cell tree induces a morphism of persistence modules, and the image, kernel, and cokernel of this morphism describe complementary aspects of organelle organization: where organelles occupy branches, where they co-localize within branch structures, and where they are absent. An efficient tree-based algorithm is developed for computing these descriptors. Applied to retinal microglia, the method reveals organelle-specific spatial programs: CD68-positive organelles reorganize in a layer- and injury-dependent manner, while mitochondrial organization remains more closely coupled to the underlying branching morphology.

The third part of the thesis studies how stable homological invariants can be used in machine learning. Stable ranks provide a bridge from persistence modules to function spaces or finite-dimensional vector spaces, making persistence-based information accessible to kernel methods and neural networks. We introduce stable rank kernels, in which the choice of distance on persistence modules determines the stable rank and, consequently, the similarities encoded by the kernel. Varying this distance through contours can improve supervised learning performance. We also study subsampling-based stable ranks, in which probability distributions on a reference dataset are used to draw many subsamples, compute persistent homology and the corresponding stable ranks, and average the resulting functions. Different choices of distribution yield global descriptors of datasets or relative descriptors of points in the ambient space with respect to a reference object.

Finally, we investigate robustness in persistence-based learning. Persistent homology is stable with respect to suitable metrics, but these guarantees need not be preserved when persistence modules are processed by neural networks. We therefore introduce a stable rank network, combining stable rank vectorizations with Lipschitz neural network layers. This architecture has a controlled Lipschitz constant and yields sample-wise certificates of robustness in Wasserstein or bottleneck distance. This shows that topological stability can be preserved through a learning pipeline and used to certify robustness against adversarial perturbations.

Abstract [sv]

Moderna datamängder blir alltmer komplexa och heterogena. För att analysera sådan data måste man först avgöra vilket matematiskt rum som är anpassat för datan, hur objekt i detta rum ska jämföras, och vilka beräkningsbara invarianter som bevarar den information som är relevant för en given uppgift. Denna avhandling studerar dessa frågor ur ett topologiskt dataanalysperspektiv, med homologi som ett språk för att beskriva datans geometri.

Den första delen av avhandlingen utvecklar avstånd och invarianter för rum av persistensmoduler. Utgångspunkten är kategoriteoretisk: vi betraktar data som objekt i kategorier med tillräckligt mycket algebraisk struktur för att definiera ett storleksbegrepp. I synnerhet har abelska kategorier kärnor, kokärnor och exakta följder, vilket gör det möjligt att konstruera avstånd mellan objekt utifrån i vilken grad morfismer misslyckas med att vara isomorfismer. Persistensmoduler utgör ett centralt exempel: de beskriver hur homologiska egenskaper uppstår och försvinner längs en filtrering, och kan ses som funktorer från ett partiellt ordnat parameterrum till vektorrum. Inom detta ramverk studerar vi metriker inducerade av contours, vilka beskriver flöden i parameterrummet. Detta leder till kompakthetsresultat för familjer av flerdimensionella persistensmoduler. I det endimensionella fallet utvecklar vi algebraiska Wasserstein-metriker genom att kombinera contours med en viktning av persistensmoduler via ℓp-normer av deras intervallängder. Dessa metriker ger upphov till Wasserstein stable ranks, vilka är stabila och beräkningsbara invarianter som kan parametriseras på ett tolkningsbart sätt och optimeras i maskininlärningsuppgifter.

Den andra delen av avhandlingen tillämpar topologiska metoder inom neurovetenskap, där cellulära morfologier ger naturliga exempel på strukturerad geometrisk data. Mikroglia och andra förgrenade celler kan representeras som rotade träd i ett tredimensionellt rum, och deras morfologi kan karakteriseras med hjälp av topologiska deskriptorer. I morphOMICs-pipelinen kombineras sådana deskriptorer med vektoriseringar, bootstrapmetoder, dimensionsreduktion och klassificering för att kartlägga mikroglians morfologi över hjärnregioner, kön, utveckling, sjukdom och experimentell perturbation. Detta ger en datadriven atlas över mikroglians morfologi som inte förlitar sig på förvalda skalära morfologiska mått.

Vi introducerar också chromatic TMD, en topologisk deskriptor för att studera intracellulär organisation i förgrenade celler. Här representeras en cell av ett rotat träd, medan organeller såsom CD68-positiva lysosomer eller mitokondrier utgör delgrafer av detta träd. Inklusionen av organell-delgrafen i cellträdet inducerar en morfism av persistensmoduler, och bilden, kärnan och kokärnan av denna morfism beskriver komplementära aspekter av organellernas organisation: var organeller förekommer i grenarna, var de samlokaliseras i grenstrukturer och var de inte förekommer. En effektiv trädbaserad algoritm utvecklas för att beräkna dessa deskriptorer. Tillämpad på mikroglia från näthinnan påvisar metoden organellspecifik struktur: CD68-positiva lysosomer uppvisar olika omorganisation beroende på näthinnelager och skada, medan den mitokondriella organisationen förblir mer tätt kopplad till den underliggande förgreningsmorfologin.

Den tredje delen av avhandlingen studerar hur stabila homologiska invarianter kan användas i maskininlärning. Stable ranks skapar en brygga mellan persistensmoduler och funktionsrum eller ändligdimensionella vektorrum, vilket gör homologisk information tillgänglig för kärnmetoder och neurala nätverk. Vi introducerar stable rank kernels, där valet av metrik på persistensmoduler avgör vilken geometri kärnan ser, och visar att variation av denna metrik genom contours kan förbättra prestanda i övervakad inlärning. Vi studerar också subsamplingsbaserade stable ranks, där man upprepade gånger samplar från en sannolikhetsfördelning på en referensdatamängd, beräknar homologiska invarianter och bildar medelvärdet av de resulterande stable ranks. Detta ger både globala deskriptorer av datamängder och relativa deskriptorer av punkter med avseende på referensobjektet.

Slutligen undersöker vi robusthet i homologibaserad inlärning. Persistent homologi är stabil med avseende på lämpliga metriker, men denna stabilitet kan gå förlorad när persistensmoduler används i obegränsade neurala nätverk. Vi introducerar därför ett stable rank network som kombinerar vektorisering av persistensmodulen med Lipschitz-neurala nätverkslager. Denna arkitektur har en kontrollerad Lipschitzkonstant och ger robusthetscertifikat per datapunkt i Wasserstein- eller bottleneck-avstånd. Resultatet visar att topologisk stabilitet kan bevaras genom en inlärningspipeline och användas för att certifiera robusthet mot adversariella perturbationer.

Place, publisher, year, edition, pages
Stockholm, Sweden: KTH Royal Institute of Technology, 2026. , p. 287
Series
TRITA-SCI-FOU ; 2026:18
Keywords [en]
topological data analysis, persistent homology
Keywords [sv]
topologisk dataanalys
National Category
Algebra and Logic
Research subject
Applied and Computational Mathematics
Identifiers
URN: urn:nbn:se:kth:diva-387758ISBN: 978-91-8106-678-4 (print)OAI: oai:DiVA.org:kth-387758DiVA, id: diva2:2097187
Public defence
2026-09-23, F3, Lindstedsvägen 26 & 28, Stockholm, 10:00 (English)
Opponent
Supervisors
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Note

QC 2026-09-01

Available from: 2026-09-01 Created: 2026-08-31 Last updated: 2026-09-01Bibliographically approved
List of papers
1. Choosing the geometry: valuations, norms, and contours
Open this publication in new window or tab >>Choosing the geometry: valuations, norms, and contours
(English)Manuscript (preprint) (Other academic)
Abstract [en]

We develop a general framework for constructing pseudometrics on objects of Abelian categories using valuations, which assign non-negative costs to morphisms. Specializing to finitely presented vector space representations of the poset [0, ∞), a natural setting for multiparameter persistence modules, we show that a large class of these distances is governed by contours, right-continuous lax actions of [0, ∞) on the parameter poset. Properties of contours translate into geometric properties of the induced distances, including conditions under which the resulting pseudometrics are metrics and allowing us to identify a rich family of compact subsets of finitely presented representations.

Keywords
persistent homology, multiparameter persistence, abelian categories
National Category
Algebra and Logic
Research subject
Mathematics
Identifiers
urn:nbn:se:kth:diva-387750 (URN)10.48550/arXiv.2608.14915 (DOI)
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Note

QC 20260831

Available from: 2026-08-28 Created: 2026-08-28 Last updated: 2026-08-31Bibliographically approved
2. Algebraic Wasserstein distances and stable homological invariants of data
Open this publication in new window or tab >>Algebraic Wasserstein distances and stable homological invariants of data
2025 (English)In: Journal of Applied and Computational Topology, ISSN 2367-1726, Vol. 9, no 1, article id 4Article in journal (Refereed) Published
Abstract [en]

Distances have a ubiquitous role in persistent homology, from the direct comparison of homological representations of data to the definition and optimization of invariants. In this article we introduce a family of parametrized pseudometrics between persistence modules based on the algebraic Wasserstein distance defined by Skraba and Turner, and phrase them in the formalism of noise systems. This is achieved by comparing p-norms of cokernels (resp. kernels) of monomorphisms (resp. epimorphisms) between persistence modules and corresponding bar-to-bar morphisms, a novel notion that allows us to bridge between algebraic and combinatorial aspects of persistence modules. We use algebraic Wasserstein distances to define invariants, called Wasserstein stable ranks, which are 1-Lipschitz stable with respect to such pseudometrics. We prove a low-rank approximation result for persistence modules which allows us to efficiently compute Wasserstein stable ranks, and we propose an efficient algorithm to compute the interleaving distance between them. Importantly, Wasserstein stable ranks depend on interpretable parameters which can be learnt in a machine learning context. Experimental results illustrate the use of Wasserstein stable ranks on real and artificial data and highlight how such pseudometrics could be useful in data analysis tasks.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
Persistence modules, Persistent homology, Stable topological invariants of data, Wasserstein metrics
National Category
Algebra and Logic
Identifiers
urn:nbn:se:kth:diva-360578 (URN)10.1007/s41468-024-00200-w (DOI)2-s2.0-85217793769 (Scopus ID)
Note

QC 20250228

Available from: 2025-02-26 Created: 2025-02-26 Last updated: 2026-08-31Bibliographically approved
3. Chromatic topological mapping reveals organelle-specific spatial organization within microglia
Open this publication in new window or tab >>Chromatic topological mapping reveals organelle-specific spatial organization within microglia
Show others...
(English)Manuscript (preprint) (Other academic)
Abstract [en]

In branched cells, including neurons and glia, intracellular organelles are distributed across complex cellular arbors where their spatial arrangement supports transport, signaling, and compartmentalized function. Although intracellular organelle organization is increasingly recognized as an important feature of cellular state and function, existing approaches assess organelle abundance or spatial position without accounting for the branching architecture that shapes cellular function. Here, we introduce the chromatic topological morphology descriptor (chromatic TMD), a framework that quantitatively resolves intracellular organization in relation to branching morphology. Applied to reconstructed microglia with annotated lysosomal and mitochondrial compartments across retinal layers and after optic nerve crush injury, chromatic TMD identifies distinct organelle-specific spatial programs: CD68+-endosomal-lysosomes undergo layer-dependent branch-specific redistribution, revealing selective intracellular reorganization after injury, whereas mitochondrial organization remains closely coupled to branching morphology. These findings establish intracellular organization as an additional layer of cellular architecture that can be systematically analyzed across branched neural cells.

Keywords
topological data analysis, microglia
National Category
Cell Biology
Research subject
Biological Physics
Identifiers
urn:nbn:se:kth:diva-387756 (URN)10.64898/2026.06.05.730307 (DOI)
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Note

QC 20260831

Available from: 2026-08-30 Created: 2026-08-30 Last updated: 2026-08-31Bibliographically approved
4. A tool for mapping microglial morphology, morphOMICs, reveals brain-region and sex-dependent phenotypes
Open this publication in new window or tab >>A tool for mapping microglial morphology, morphOMICs, reveals brain-region and sex-dependent phenotypes
Show others...
2022 (English)In: Nature Neuroscience, ISSN 1097-6256, E-ISSN 1546-1726, Vol. 25, no 10, p. 1379-+Article in journal (Refereed) Published
Abstract [en]

Environmental cues influence the highly dynamic morphology of microglia. Strategies to characterize these changes usually involve user-selected morphometric features, which preclude the identification of a spectrum of context-dependent morphological phenotypes. Here we develop MorphOMICs, a topological data analysis approach, which enables semiautomatic mapping of microglial morphology into an atlas of cue-dependent phenotypes and overcomes feature-selection biases and biological variability. We extract spatially heterogeneous and sexually dimorphic morphological phenotypes for seven adult mouse brain regions. This sex-specific phenotype declines with maturation but increases over the disease trajectories in two neurodegeneration mouse models, with females showing a faster morphological shift in affected brain regions. Remarkably, microglia morphologies reflect an adaptation upon repeated exposure to ketamine anesthesia and do not recover to control morphologies. Finally, we demonstrate that both long primary processes and short terminal processes provide distinct insights to morphological phenotypes. MorphOMICs opens a new perspective to characterize microglial morphology.

Place, publisher, year, edition, pages
Springer Nature, 2022
National Category
Cell Biology Biophysics
Identifiers
urn:nbn:se:kth:diva-320482 (URN)10.1038/s41593-022-01167-6 (DOI)000862214700001 ()36180790 (PubMedID)2-s2.0-85139248488 (Scopus ID)
Note

QC 20221026

Available from: 2022-10-26 Created: 2022-10-26 Last updated: 2026-08-31Bibliographically approved
5. Certifying Robustness via Topological Representations
Open this publication in new window or tab >>Certifying Robustness via Topological Representations
Show others...
(English)Manuscript (preprint) (Other academic)
Abstract [en]

Exploring the shape of data spaces is providing new insights in data analysis and deep learning tasks within a variety of application domains. A common approach in Topological Data Analysis to extract multi-scale intrinsic geometric properties of data is persistent homology. This method enjoys theoretical stability results (i.e. Lipschitz continuity with respect to appropriate metrics), however the significance of this robustness when persistent homology is used in machine learning is underexplored. We propose a neural network architecture that can learn discriminative geometric representations from persistence with a controllable Lipschitz constant. In adversarial learning, this end-to-end stability can be used to certify ε-robustness for samples in a dataset, which we demonstrate on the ORBIT5K data set representing the orbits of a discrete dynamical system.

Keywords
topological data analysis, adversarial machine learning
National Category
Algebra and Logic
Research subject
Mathematics
Identifiers
urn:nbn:se:kth:diva-387757 (URN)
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Note

QC 20260831

Available from: 2026-08-30 Created: 2026-08-30 Last updated: 2026-08-31Bibliographically approved
6. Supervised Learning Using Homology Stable Rank Kernels
Open this publication in new window or tab >>Supervised Learning Using Homology Stable Rank Kernels
2021 (English)In: FRONTIERS IN APPLIED MATHEMATICS AND STATISTICS, ISSN 2297-4687, Vol. 7, article id 668046Article in journal (Refereed) Published
Abstract [en]

Exciting recent developments in Topological Data Analysis have aimed at combining homology-based invariants with Machine Learning. In this article, we use hierarchical stabilization to bridge between persistence and kernel-based methods by introducing the so-called stable rank kernels. A fundamental property of the stable rank kernels is that they depend on metrics to compare persistence modules. We illustrate their use on artificial and real-world datasets and show that by varying the metric we can improve accuracy in classification tasks.

Place, publisher, year, edition, pages
FRONTIERS MEDIA SA, 2021
Keywords
topological data analysis, kernel methods, metrics, hierarchical stabilisation, persistent homology
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-299493 (URN)10.3389/fams.2021.668046 (DOI)000677390900001 ()2-s2.0-85111102378 (Scopus ID)
Note

QC 20210809

Available from: 2021-08-09 Created: 2021-08-09 Last updated: 2026-08-31Bibliographically approved
7. Global and Relative Topological Features from Homological Invariants of Subsampled Datasets
Open this publication in new window or tab >>Global and Relative Topological Features from Homological Invariants of Subsampled Datasets
2023 (English)In: Proceedings of the 2nd Annual Topology, Algebra, and Geometry in Machine Learning, TAG-ML 2023, ML Research Press , 2023, p. 302-312Conference paper, Published paper (Refereed)
Abstract [en]

Homology-based invariants can be used to characterize the geometry of datasets and thereby gain some understanding of the processes generating those datasets. In this work we investigate how the geometry of a dataset changes when it is subsampled in various ways. In our framework the dataset serves as a reference object; we then consider different points in the ambient space and endow them with a geometry defined in relation to the reference object, for instance by subsampling the dataset proportionally to the distance between its elements and the point under consideration. We illustrate how this process can be used to extract rich geometrical information, allowing for example to classify points coming from different data distributions.

Place, publisher, year, edition, pages
ML Research Press, 2023
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-340790 (URN)001220893300023 ()2-s2.0-85178663624 (Scopus ID)
Conference
2nd Annual Workshop on Topology, Algebra, and Geometry in Machine Learning, TAG-ML 2023, held at the International Conference on Machine Learning, ICML 2023, Honolulu, United States of America, Jul 28 2023
Note

QC 20231215

Available from: 2023-12-15 Created: 2023-12-15 Last updated: 2026-08-31Bibliographically approved

Open Access in DiVA

fulltext(2311 kB)29 downloads
File information
File name FULLTEXT01.pdfFile size 2311 kBChecksum SHA-512
054db487c8997c4c7b2faf5fbbbb53ae5a8557a21cbdab81e52a32bd5ebaf065e702d24d68bdfc6aca0321ae7a1e1e17908adaf15da087d854db73898c720f8a
Type fulltextMimetype application/pdf

Authority records

Agerberg, Jens

Search in DiVA

By author/editor
Agerberg, Jens
By organisation
Algebra, Combinatorics and Topology
Algebra and Logic

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

isbn
urn-nbn

Altmetric score

isbn
urn-nbn
Total: 646 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf