kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Publications (3 of 3) Show all publications
Segeljakt, K., Haridi, S. & Carbone, P. (2024). AquaLang: A Dataflow Programming Language. In: DEBS 2024 - Proceedings of the 18th ACM International Conference on Distributed and Event-Based Systems: . Paper presented at 18th ACM International Conference on Distributed and Event-Based Systems, DEBS 2024, Villeurbanne, France, Jun 25 2024 - Jun 28 2024 (pp. 42-53). Association for Computing Machinery (ACM)
Open this publication in new window or tab >>AquaLang: A Dataflow Programming Language
2024 (English)In: DEBS 2024 - Proceedings of the 18th ACM International Conference on Distributed and Event-Based Systems, Association for Computing Machinery (ACM) , 2024, p. 42-53Conference paper, Published paper (Refereed)
Abstract [en]

Dataflow systems are widely used today for building and running continuous data-intensive applications. However, the unavoidable semantic gap between the host languages of dataflow system libraries and the dataflow model creates programmability limitations that hinder performance, safety, and ease of use. We propose AquaLang, a new language designed for dataflow systems. Programs in AquaLang blend strongly typed relational and functional syntax and are verified using an effect system that prevents undefined behaviour that can occur when introducing user-defined logic that violates dataflow semantics. Unverified external code is also feasible in AquaLang through the novel use of sandboxing. Furthermore, on top of standard dataflow optimisations employed by current systems, AquaLang's ability to analyze algebraic properties of user-defined functions further unlocks the potential of deeper dataflow program re-writing. In our evaluation, we measure up to one order of magnitude speedup for Nexmark queries against hand-written Flink programs attributed to pushdown and window incrementalisation techniques.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2024
Keywords
Data Streams, Dataflow Systems, Programming Languages
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-351925 (URN)10.1145/3629104.3666030 (DOI)001283849100006 ()2-s2.0-85200659561 (Scopus ID)
Conference
18th ACM International Conference on Distributed and Event-Based Systems, DEBS 2024, Villeurbanne, France, Jun 25 2024 - Jun 28 2024
Note

Part of ISBN 9798400704437

QC 20240827

Available from: 2024-08-19 Created: 2024-08-19 Last updated: 2024-09-10Bibliographically approved
Kroll, L., Segeljakt, K., Schulte, C., Haridi, S. & Carbone, P. (2019). Arc: An IR for batch and stream programming. In: Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI): . Paper presented at 17th ACM SIGPLAN International Symposium on Database Programming Languages, DBPL 2019, co-located with PLDI 2019; Phoenix; United States; 23 June 2019 (pp. 53-58). Association for Computing Machinery (ACM)
Open this publication in new window or tab >>Arc: An IR for batch and stream programming
Show others...
2019 (English)In: Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), Association for Computing Machinery (ACM), 2019, p. 53-58Conference paper, Published paper (Refereed)
Abstract [en]

In big data analytics, there is currently a large number of data programming models and their respective frontends such as relational tables, graphs, tensors, and streams. This has lead to a plethora of runtimes that typically focus on the efficient execution of just a single frontend. This fragmentation manifests itself today by highly complex pipelines that bundle multiple runtimes to support the necessary models. Hence, joint optimization and execution of such pipelines across these frontend-bound runtimes is infeasible. We propose Arc as the first unified Intermediate Representation (IR) for data analytics that incorporates stream semantics based on a modern specification of streams, windows and stream aggregation, to combine batch and stream computation models. Arc extends Weld, an IR for batch computation and adds support for partitioned, out-of-order stream and window operators which are the most fundamental building blocks in contemporary data streaming.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2019
Series
Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI)
Keywords
Data analytics, Intermediate representation, Stream processing
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:kth:diva-262617 (URN)10.1145/3315507.3330199 (DOI)000519102100007 ()2-s2.0-85071162040 (Scopus ID)
Conference
17th ACM SIGPLAN International Symposium on Database Programming Languages, DBPL 2019, co-located with PLDI 2019; Phoenix; United States; 23 June 2019
Note

QC 20191017

Part of ISBN 9781450367189

Available from: 2019-10-17 Created: 2019-10-17 Last updated: 2024-10-15Bibliographically approved
Meldrum, M., Segeljakt, K., Kroll, L., Carbone, P., Schulte, C. & Haridi, S. (2019). Arcon: Continuous and deep data stream analytics. In: ACM International Conference Proceeding Series: . Paper presented at 13th International Workshop on Real-Time Business Intelligence and Analytics, BIRTE 2019, in conjunction with the VLDB 2019 Conference, 26 August 2019. Association for Computing Machinery
Open this publication in new window or tab >>Arcon: Continuous and deep data stream analytics
Show others...
2019 (English)In: ACM International Conference Proceeding Series, Association for Computing Machinery , 2019Conference paper, Published paper (Refereed)
Abstract [en]

Contemporary end-to-end data pipelines need to combine many diverse workloads such as machine learning, relational operations, stream dataflows, tensor transformations, and graphs. For each of these workload types, there exists several frontends (e.g., SQL, Beam, Keras) based on different programming languages as well as different runtimes (e.g., Spark, Flink, Tensorflow) that optimize for a particular frontend and possibly a hardware architecture (e.g., GPUs). The resulting pipelines suffer in terms of complexity and performance due to excessive type conversions, materialization of intermediate results, and lack of cross-framework optimizations. Arcon aims to provide a unified approach to declare and execute tasks across frontend-boundaries as well as enabling their seamless integration with event-driven services at scale. In this demonstration, we present Arcon and through a series of use-case scenarios demonstrate that its execution model is powerful enough to cover existing as well as upcoming real-time computations for analytics and application-specific needs.

Place, publisher, year, edition, pages
Association for Computing Machinery, 2019
Keywords
Data flow analysis, Information analysis, Object oriented programming, Program processors, Application specific, Framework optimization, Hardware architecture, Intermediate results, Real-time computations, Relational operations, Seamless integration, Tensor transformation, Pipelines
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:kth:diva-268550 (URN)10.1145/3350489.3350492 (DOI)2-s2.0-85072806432 (Scopus ID)
Conference
13th International Workshop on Real-Time Business Intelligence and Analytics, BIRTE 2019, in conjunction with the VLDB 2019 Conference, 26 August 2019
Note

QC 20200324

Part of ISBN 9781450376600

Available from: 2020-03-24 Created: 2020-03-24 Last updated: 2024-10-15Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-7096-4401

Search in DiVA

Show all publications