Foundations of Trustworthy AI-Native Data Systems
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]
In traditional data management systems, queries have well-defined semantics and produce exact results. Integrating Machine Learning inference into data processing pipelines disrupts both properties by introducing operators whose outputs are approximate rather than exact. This thesis establishes two foundations for trustworthy AI-native data systems: empirical characterization of ML operator execution cost, and formal, declarative correctness guarantees that the system enforces on behalf of the user. We develop these foundations across three levels of abstraction, from single-operator cost, to single-operator correctness, to their joint optimization at the pipeline level, and establish Conformal Prediction as a practical statistical foundation for this approach. We introduce Crayfish, a benchmarking framework for ML inference within dataflow engines that reveals how interactions between serving tools, stream processors, and pipeline configurations shape inference costs in ways that are difficult to anticipate from component-level behavior alone. We propose ConANN, the first framework to provide distribution-free recall guarantees for Inverted File-based Approximate Nearest Neighbor search, using conformal methods to replace heuristic index tuning with formal statistical guarantees. At the pipeline level, we study joint cost and correctness optimization in the context of Neural Graph Databases, where multi-hop queries over Knowledge Graphs interleave retrieval and neural execution. We formalize a hybrid query optimization architecture for this setting, then introduce ConRAD, which enforces end-to-end recall guarantees for multi-hop queries while dynamically bypassing expensive neural inference when recall targets can be met with local graph evidence alone. Taken together, these contributions show that the rigor users expect from traditional data systems need not be abandoned as those systems become increasingly driven by Machine Learning.
Abstract [sv]
I traditionella datahanteringssystem har frågor väldefinierad semantik och producerar exakta resultat. Att integrera maskininlärningsinferens i databehandlingspipelines stör båda dessa egenskaper genom att introducera operatorer vars utdata är approximativa snarare än exakta. Denna avhandling etablerar två grundpelare för tillförlitliga AI-nativa datasystem: empirisk karaktärisering av exekveringskostnaden för ML-operatorer, samt formella, deklarativa korrekthetsgarantier som systemet upprätthåller å användarens vägnar. Vi utvecklar dessa grundpelare över tre abstraktionsnivåer, från enskild operatorkostnad till enskild operatorkorrekthet, och slutligen till deras gemensamma optimering på pipelinenivå. Vi etablerar Conformal Prediction som en praktisk statistisk grund för detta tillvägagångssätt. Vi introducerar Crayfish, ett benchmarkingramverk för ML-inferens inom dataflödesmotorer som synliggör hur interaktioner mellan serving-verktyg, strömprocessorer och pipelinekonfigurationer formar inferenskostnaden på sätt som är svåra att förutse enbart utifrån enskilda komponenters beteende. Vi föreslår ConANN, det första ramverket som erbjuder distributionsfria recall-garantier för Inverted File-baserad approximativ närmaste-granne-sökning, genom att använda konforma metoder för att ersätta heuristisk indexjustering med formella statistiska garantier. På pipelinenivå studerar vi gemensam optimering av kostnad och korrekthet i kontexten av neurala grafdatabaser, där flerstegsfrågor över kunskapsgrafer varvar hämtning med neural exekvering. Vi formaliserar en hybrid frågeoptimeringsarkitektur för detta scenario och introducerar sedan ConRAD, som upprätthåller end-to-end recall-garantier för flerstegsfrågor och samtidigt dynamiskt kringgår kostsam neural inferens när recall-målen kan uppnås med enbart lokal grafevidens. Sammantaget visar dessa bidrag att den stringens som användare förväntar sig av traditionella datasystem inte behöver överges i takt med att dessa system i allt högre grad drivs av maskininlärning.
Place, publisher, year, edition, pages
Stockholm, Sweden: KTH Royal Institute of Technology, 2026. , p. 73
Keywords [en]
AI-Native Data Systems, Model Serving, Conformal Risk Control, Approximate Nearest Neighbor Search, Neural Graph Databases, Query Optimization
Keywords [sv]
AI-nativa datasystem, Modellservering, Konform riskkontroll, Approximativ närmaste-granne-sökning, Neurala grafdatabaser, Frågeoptimering
National Category
Computer Systems
Identifiers
URN: urn:nbn:se:kth:diva-382126ISBN: 978-91-8106-628-9 (print)OAI: oai:DiVA.org:kth-382126DiVA, id: diva2:2061824
Public defence
2026-06-15, https://kth-se.zoom.us/j/65502477126, F3 Flodis, Lindstedtvägen 26, Stockholm, 14:00 (English)
Opponent
Supervisors
Note
QC 20260522
2026-05-222026-05-222026-06-29Bibliographically approved
List of papers