kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Publications (7 of 7) Show all publications
Hakimzadeh, K., Nicholsonz, P. K., Lugonesz, D. & Payberah, A. H. (2021). IMITA: Imitation Learning for Generalizing Cloud Orchestration. In: Lefevre, L Patterson, S Lee, YC Shen, H Ilager, S Goudarzi, M Toosi, AN Buyya, R (Ed.), 21St IEEE/ACM International Symposium On Cluster, Cloud And Internet Computing (CCGRID 2021): . Paper presented at 21st IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid), MAY 10-13, 2021, ELECTR NETWORK (pp. 237-246). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>IMITA: Imitation Learning for Generalizing Cloud Orchestration
2021 (English)In: 21St IEEE/ACM International Symposium On Cluster, Cloud And Internet Computing (CCGRID 2021) / [ed] Lefevre, L Patterson, S Lee, YC Shen, H Ilager, S Goudarzi, M Toosi, AN Buyya, R, Institute of Electrical and Electronics Engineers (IEEE) , 2021, p. 237-246Conference paper, Published paper (Refereed)
Abstract [en]

Operating large scale and feature-rich applications is becoming increasingly complex as engineers need to deploy highly configurable software releases on distributed cloud stacks while managing ever-shorter production cycles. Although recent proposals attempt to streamline cloud resources orchestration, there is still a significant challenge in making such solutions generalize to unseen cloud stacks. In other words, the behavior of application-specific Key Performance Indicators (KPIs) and resource configurations, crafted for specific stacks, may differ on heterogeneous deployments, requiring time-consuming policy adjustments. We introduce IMITA, a system that leverages imitation learning to create models by imitating an expert behavior that can be generalized seamlessly to new cloud stacks. To make a generalized model, IMITA maps expert actions taken based on the application KPI space to the space of resource utilization metrics that are universally available in cloud platforms. This mapping enables the model to trigger actions, mimicking expert behavior, upon the occurrence of similar resource utilization footprints across deployments. We demonstrate IMITA by learning to scale-out Cassandra deployments with diverse configurations and workloads. Our results show IMITA can replicate expert actions across deployments and extrapolate to unseen environments by achieving 50 - 94% fewer false positives actions than traditional threshold-based policies while still adhering to Service-Level Objectives (SLO) and avoiding under-provisioning of resources. Moreover, since collecting data in clouds is costly, IMITA gathers data only for representative configurations to train the imitator model. This approach reduces the size of the collected data to 50%.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2021
Keywords
Cloud Orchestration, Imitation Learning, Generalization
National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-304178 (URN)10.1109/CCGrid51090.2021.00033 (DOI)000703983200024 ()2-s2.0-85114885031 (Scopus ID)
Conference
21st IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid), MAY 10-13, 2021, ELECTR NETWORK
Note

Part of proceedings: ISBN 978-1-7281-9586-5, QC 20230117

Available from: 2021-11-05 Created: 2021-11-05 Last updated: 2023-01-17Bibliographically approved
Hakimzadeh, K. & Dowling, J. (2019). Karamel: A System for Timely Provisioning Large-Scale Software Across IaaS Clouds. In: 2019 IEEE 12th International Conference on Cloud Computing (CLOUD): . Paper presented at 12th IEEE International Conference on Cloud Computing, CLOUD 2019, Milan, Italy, July 8-13, 2019 (pp. 391-395). IEEE Computer Society, Article ID 8814511.
Open this publication in new window or tab >>Karamel: A System for Timely Provisioning Large-Scale Software Across IaaS Clouds
2019 (English)In: 2019 IEEE 12th International Conference on Cloud Computing (CLOUD), IEEE Computer Society, 2019, p. 391-395, article id 8814511Conference paper, Published paper (Refereed)
Abstract [en]

Cloud-native systems and application software platforms are becoming increasingly complex, and, ideally, they are expected to be quick to launch, elastic, portable across different cloud environments and easily managed. However, as cloud applications increase in complexity, so do the resultant challenges in configuring such applications, and orchestrating the deployment of their constituent services on potentially different cloud operating systems and environments.

This paper presents a new orchestration system called Karamel that addresses these challenges by providing a cloud-independent orchestration service for deploying and configuring cloud applications and platforms across different environments. In Karamel, we model configuration routines with their dependencies in composable modules, and we achieve a high level of configuration/deployment parallelism by using techniques such as DAG traversal control logic, dataflow variable binding, and parallel actuation. In Karamel, complex distributed deployments are specified declaratively in a compact YAML syntax, and cluster definitions can be validated using an external artifact repository (GitHub).

Place, publisher, year, edition, pages
IEEE Computer Society, 2019
Series
IEEE International Conference on Cloud Computing, CLOUD, ISSN 2159-6182
Keywords
cloud computing, distributed systems, configuration management, orchestration
National Category
Computer Sciences
Research subject
Computer Science
Identifiers
urn:nbn:se:kth:diva-254525 (URN)10.1109/CLOUD.2019.00069 (DOI)000556208000056 ()2-s2.0-85072313095 (Scopus ID)
Conference
12th IEEE International Conference on Cloud Computing, CLOUD 2019, Milan, Italy, July 8-13, 2019
Note

QC 20190710

Available from: 2019-07-01 Created: 2019-07-01 Last updated: 2022-06-26Bibliographically approved
Hakimzadeh, K. & Dowling, J. (2019). Ops-Scale: Scalable and Elastic Cloud Operations by a Functional Abstraction and Feedback Loops. In: 2019 IEEE 13th International Conference on Self-Adaptive And Self-Organizing Systems (SASO): . Paper presented at 13th IEEE International Conference on Self-Adaptive and Self-Organizing Systems, SASO 2019; Umeå; Sweden; 16 June 2019 through 20 June 2019 (pp. 62-71). IEEE Computer Society, Article ID 8780537.
Open this publication in new window or tab >>Ops-Scale: Scalable and Elastic Cloud Operations by a Functional Abstraction and Feedback Loops
2019 (English)In: 2019 IEEE 13th International Conference on Self-Adaptive And Self-Organizing Systems (SASO), IEEE Computer Society, 2019, p. 62-71, article id 8780537Conference paper, Published paper (Refereed)
Abstract [en]

Recent research has proposed new techniques to streamline the autoscaling of cloud applications, but little effort has been made to advance configuration management (CM) systems for such elastic operations. Existing practices use CM systems, from the DevOps paradigm, to automate operations. However, these practices still require human intervention to program ad hoc procedures to fully automate reconfiguration. Moreover, even after careful programming of cloud operations, the backing models are insufficient for re-running such programs unchanged in other platforms - which implies an overhead in rewriting the programs. We argue that CM programs can be designed to be deployment-agnostic and highly elastic with well-defined abstractions. In this paper, we introduce our abstraction based on declarative functional programming, and we demonstrate it using a feedback loop control mechanism. Our proposal, called Ops-Scale, is a family of cloud operations that are derived by making a functional abstraction over existing configuration programs. The hypothesis in this paper is twofold: 1) it should be possible to make a highly declarative CM system rich enough to capture fine-grained reconfigurations of autoscaling automatically, and; 2) that a program written for a specific deployment can be re-used in other deployments. To test this hypothesis, we have implemented an open source configuration engine called Karamel that is already used in industry for large-scale cluster deployments. Results show that at scale Ops-Scale can capture a polynomial order of reconfiguration growth in a fully automated manner. In practice, recent deployments have demonstrated that Karamel can provision clusters of 100 virtual machines consisting of many-layers distributed services on Google's IaaS Cloud in 'less than 10 minutes'.

Place, publisher, year, edition, pages
IEEE Computer Society, 2019
Series
International Conference on Self-Adaptive and Self-Organizing Systems, ISSN 1949-3673
Keywords
Cloud Computing, Functional Programming, Elasticity, Auto-Scaling, Feedback Control Loop
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-254521 (URN)10.1109/SASO.2019.00017 (DOI)000501005500008 ()2-s2.0-85070557782 (Scopus ID)
Conference
13th IEEE International Conference on Self-Adaptive and Self-Organizing Systems, SASO 2019; Umeå; Sweden; 16 June 2019 through 20 June 2019
Note

QC 20200114

Part of ISBN 978-1-7281-2731-6

Available from: 2019-07-01 Created: 2019-07-01 Last updated: 2024-10-21Bibliographically approved
Hakimzadeh, K., Nicholson, P. K. & Lugones, D. (2018). Auto-scaling with apprenticeship learning. In: SoCC 2018 - Proceedings of the 2018 ACM Symposium on Cloud Computing: . Paper presented at 2018 ACM Symposium on Cloud Computing, SoCC 2018, Carlsbad, United States, 11 October 2018 through 13 October 2018 (pp. 512-512). Association for Computing Machinery (ACM)
Open this publication in new window or tab >>Auto-scaling with apprenticeship learning
2018 (English)In: SoCC 2018 - Proceedings of the 2018 ACM Symposium on Cloud Computing, Association for Computing Machinery (ACM), 2018, p. 512-512Conference paper, Published paper (Refereed)
Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2018
National Category
Other Engineering and Technologies
Identifiers
urn:nbn:se:kth:diva-241482 (URN)10.1145/3267809.3275454 (DOI)000458692200049 ()2-s2.0-85059007117 (Scopus ID)9781450360111 (ISBN)
Conference
2018 ACM Symposium on Cloud Computing, SoCC 2018, Carlsbad, United States, 11 October 2018 through 13 October 2018
Note

QC 20190123

Available from: 2019-01-23 Created: 2019-01-23 Last updated: 2022-06-26Bibliographically approved
Peiro Sajjad, H., Hakimzadeh Harirbaf, K. & Perera, S. (2017). Reproducible Distributed Clusters with Mutable Containers: To Minimize Cost and Provisioning Time. In: HotConNet '17 Proceedings of the Workshop on Hot Topics in Container Networking and Networked Systems: . Paper presented at ACM SIGCOMM 2017 1st International Workshop on Hot Topics in Container Networking and Networked Systems (HotConNet’17), Los Angeles, CA, USA on August 21-25, 2017. (pp. 18-23). Association for Computing Machinery (ACM)
Open this publication in new window or tab >>Reproducible Distributed Clusters with Mutable Containers: To Minimize Cost and Provisioning Time
2017 (English)In: HotConNet '17 Proceedings of the Workshop on Hot Topics in Container Networking and Networked Systems, Association for Computing Machinery (ACM), 2017, p. 18-23Conference paper, Published paper (Refereed)
Abstract [en]

Reproducible and repeatable provisioning of large-scale distributed systems is laborious. The cost of virtual infrastructure and the provisioning complexity are two of the main concerns. The trade-offs between virtual machines (VMs) and containers, the most popular virtualization technologies, further complicate the problem. Although containers incur little overhead compared to VMs, VMs are required for their certain guarantees such as hardware isolation.

In this paper, we present a mutable container provisioning solution, enabling users to switch infrastructure between VMs and containers seamlessly. Our solution allows for significant infrastructure-cost optimizations. We discuss that immutable containers come short for certain provisioning scenarios. However, mutable containers can incur a large time overhead. To reduce the time overhead, we propose multiple provisioning-time optimizations. We implement our solution in Karamel, an open-sourced reproducible provisioning system. Based on our evaluation results, we discuss the cost-optimization opportunities and the time-optimization challenges of our new model.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2017
Keywords
Containers, Reproducible Clusters, Mutable, Provisioning, Cloud
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:kth:diva-213398 (URN)10.1145/3094405.3094409 (DOI)000426563600004 ()2-s2.0-85030776618 (Scopus ID)9781450350587 (ISBN)
Conference
ACM SIGCOMM 2017 1st International Workshop on Hot Topics in Container Networking and Networked Systems (HotConNet’17), Los Angeles, CA, USA on August 21-25, 2017.
Note

QC 20170831

Available from: 2017-08-30 Created: 2017-08-30 Last updated: 2024-03-18Bibliographically approved
Bessani, A., Brandt, J., Bux, M., Cogo, V., Dimitrova, L., Dowling, J., . . . Zimmermann, K. (2016). BiobankCloud: A platform for the secure storage, sharing, and processing of large biomedical data sets. In: 1st International Workshop on Data Management and Analytics for Medicine and Healthcare, DMAH 2015 and Workshop on Big-Graphs Online Querying, Big-O(Q) 2015 held in conjunction with 41st International Conference on Very Large Data Bases, VLDB 2015: . Paper presented at 31 August 2015 through 4 September 2015 (pp. 89-105). Springer
Open this publication in new window or tab >>BiobankCloud: A platform for the secure storage, sharing, and processing of large biomedical data sets
Show others...
2016 (English)In: 1st International Workshop on Data Management and Analytics for Medicine and Healthcare, DMAH 2015 and Workshop on Big-Graphs Online Querying, Big-O(Q) 2015 held in conjunction with 41st International Conference on Very Large Data Bases, VLDB 2015, Springer, 2016, p. 89-105Conference paper, Published paper (Refereed)
Abstract [en]

Biobanks store and catalog human biological material that is increasingly being digitized using next-generation sequencing (NGS). There is, however, a computational bottleneck, as existing software systems are not scalable and secure enough to store and process the incoming wave of genomic data from NGS machines. In the BiobankCloud project, we are building a Hadoop-based platform for the secure storage, sharing, and parallel processing of genomic data. We extended Hadoop to include support for multi-tenant studies, reduced storage requirements with erasure coding, and added support for extensible and consistent metadata. On top of Hadoop, we built a scalable scientific workflow engine featuring a proper workflow definition language focusing on simple integration and chaining of existing tools, adaptive scheduling on Apache Yarn, and support for iterative dataflows. Our platform also supports the secure sharing of data across different, distributed Hadoop clusters. The software is easily installed and comes with a user-friendly web interface for running, managing, and accessing data sets behind a secure 2-factor authentication. Initial tests have shown that the engine scales well to dozens of nodes. The entire system is open-source and includes pre-defined workflows for popular tasks in biomedical data analysis, such as variant identification, differential transcriptome analysis using RNA-Seq, and analysis of miRNA-Seq and ChIP-Seq data.

Place, publisher, year, edition, pages
Springer, 2016
Keywords
Biological materials, Data handling, Engines, Genes, Information management, Open source software, Open systems, RNA, Storage (materials), Adaptive scheduling, Biomedical data analysis, Computational bottlenecks, Next-generation sequencing, Parallel processing, Scientific workflow engines, Storage requirements, Transcriptome analysis, Digital storage
National Category
Bio Materials Other Computer and Information Science
Identifiers
urn:nbn:se:kth:diva-195513 (URN)10.1007/978-3-319-41576-5_7 (DOI)000387957300007 ()2-s2.0-84977528973 (Scopus ID)9783319415758 (ISBN)
Conference
31 August 2015 through 4 September 2015
Note

QC 20161110

Available from: 2016-11-10 Created: 2016-11-03 Last updated: 2024-03-18Bibliographically approved
Bux, M., Brandt, J., Lipka, C., Hakimzadeh, K., Dowling, J. & Leser, U. (2015). SAASFEE: Scalable scientific workflow execution engine. Paper presented at 11 September 2006 through 11 September 2006, Seoul. Proceedings of the VLDB Endowment, 8(12), 1892-1895
Open this publication in new window or tab >>SAASFEE: Scalable scientific workflow execution engine
Show others...
2015 (English)In: Proceedings of the VLDB Endowment, E-ISSN 2150-8097, Vol. 8, no 12, p. 1892-1895Article in journal (Refereed) Published
Abstract [en]

Across many fields of science, primary data sets like sensor read-outs, time series, and genomic sequences are analyzed by complex chains of specialized tools and scripts exchanging intermediate results in domain-specific file formats. Scientific work ow management systems (SWfMSs) support the development and execution of these tool chains by providing work ow specification languages, graphical editors, fault-tolerant execution engines, etc. However, many SWfMSs are not prepared to handle large data sets because of inadequate support for distributed computing. On the other hand, most SWfMSs that do support distributed computing only allow static task execution orders. We present SAASFEE, a SWfMS which runs arbitrarily complex work ows on Hadoop YARN. Work ows are specified in Cuneiform, a functional work ow language focusing on parallelization and easy integration of existing software. Cuneiform work ows are executed on Hi-WAY, a higher-level scheduler for running work ows on YARN. Distinct features of SAASFEE are the ability to execute iterative work ows, an adaptive task scheduler, re-executable provenance traces, and compatibility to selected other work ow systems. In the demonstration, we present all components of SAASFEE using real-life work ows from the field of genomics.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2015
Keywords
Chains, Computational linguistics, Engines, Information management, Specification languages, Wool, Yarn, Execution engine, Genomic sequence, Graphical editors, Intermediate results, Management systems, Parallelizations, Scientific workflows, Specialized tools, Distributed computer systems
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-181254 (URN)10.14778/2824032.2824094 (DOI)2-s2.0-84953879839 (Scopus ID)
Conference
11 September 2006 through 11 September 2006, Seoul
Note

QC 20160205

Available from: 2016-02-05 Created: 2016-01-29 Last updated: 2024-01-16Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-7781-3104

Search in DiVA

Show all publications