kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Alternative names
Publications (10 of 15) Show all publications
Wilhelmsson, M., Ismail, M. & Warsame, A. (2022). Gentrification effects on housing prices in neighbouring areas. International Journal of Housing Markets and Analysis, 15(4), 910-929
Open this publication in new window or tab >>Gentrification effects on housing prices in neighbouring areas
2022 (English)In: International Journal of Housing Markets and Analysis, ISSN 1753-8270, E-ISSN 1753-8289, Vol. 15, no 4, p. 910-929Article in journal (Refereed) Published
Abstract [en]

Purpose: This study aims to measure the occurrence of gentrification and to relate gentrification with housing values. Design/methodology/approach: The authors have used Getis-Ord statistics to identify and quantify gentrification in different residential areas in a case study of Stockholm, Sweden. Gentrification will be measured in two dimensions, namely, income and population. In step two, this measure is included in a traditional hedonic pricing model where the intention is to explain future housing prices. Findings: The results indicate that the parameter estimate is statistically significant, suggesting that gentrification contributes to higher housing values in gentrified areas and near gentrified neighbourhoods. This latter possible spillover effect of house prices due to gentrification by income and population was similar in both the hedonic price and treatment effect models. According to the hedonic price model, proximity to the gentrified area increases housing value by around 6%–8%. The spillover effect on price distribution seems to be consistent and stable in gentrified areas. Originality/value: A few studies estimate the effect of gentrification on property values. Those studies focussed on analysing the impacts of gentrification in higher rents and increasing house prices within the gentrifying areas, not gentrification on property prices in neighbouring areas. Hence, one of the paper’s contributions is to bridge the gap in previous studies by measuring gentrification’s impact on neighbouring housing prices. 

Place, publisher, year, edition, pages
Emerald, 2022
Keywords
gentrification, Getis-Ord statistics, Housing market analysis, Housing prices, spillover effects, Sweden
National Category
Economics
Identifiers
urn:nbn:se:kth:diva-311081 (URN)10.1108/IJHMA-04-2021-0049 (DOI)000685290700001 ()2-s2.0-85112584371 (Scopus ID)
Note

QC 20250331

Available from: 2022-04-19 Created: 2022-04-19 Last updated: 2025-03-31Bibliographically approved
Ismail, M. (2020). Distributed File System Metadata and its Applications. (Doctoral dissertation). KTH Royal Institute of Technology
Open this publication in new window or tab >>Distributed File System Metadata and its Applications
2020 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Distributed hierarchical file systems typically decouple the storage and serving of the file metadata from the file contents (file system blocks) to enable the file system to scale to store more data and support higher throughput. We designed HopsFS to take the scalability of the file system one step further by also decoupling the storage and serving of the file system metadata. HopsFS is an open-source, next- generation distribution of the Apache Hadoop Distributed File System (HDFS) that replaces the main scalability bottleneck in HDFS, the single-node in-memory metadata service, with a distributed metadata service built on a NewSQL database (NDB). HopsFS stores the file system’s metadata fully normalized in NDB, then it uses locking primitives and application-defined locks to ensure strongly consistent metadata.In this thesis, we leverage the consistent distributed hierarchical file system meta- data provided by HopsFS to efficiently build new classes of applications that are tightly coupled with the file system as well as to improve the internal file system operations. First, we introduce hbr, a new block reporting protocol for HopsFS that removes a scalability bottleneck that prevented HopsFS from scaling to tens of thousands of servers. Second, we introduce HopsFS-CL, a highly available cloud-native distribution of HopsFS that deploys the file system across Availability Zones in the cloud while maintaining the same file system semantics. Third, we introduce HopsFS-S3, a highly available cloud-native distribution of HopsFS that uses object stores as a backend for the block storage layer in the cloud while again maintaining the same file system semantics. Fourth, we introduce ePipe, a databus that both creates a consistent change stream for HopsFS and eventually delivers the correctly ordered stream with low latency to downstream clients. That is, ePipe extends HopsFS with a change-data-capture (CDC) API that provides not only efficient file system notifications but also enables polyglot storage for file system metadata. Polyglot storage enables us to offload metadata queries to a more appropriate engine - we use Elasticsearch to provide a free-text search of the file system namespace to demonstrate this capability. Finally, we introduce Hopsworks, a scalable, project-based multi-tenant big data platform that provides support for collaborative development and operations for teams through extended metadata.

Abstract [sv]

Distribuerade hierarkiska filsystem kopplar vanligtvis bort lagring och hanteringen av filens metadata från filens innehåll (filsystemets block) för att göra det möjligt för filsystemet att skala bättre för att lagra mer data och stödja högre genomströmn- ing. Vi utformade HopsFS för att ta skalbarheten i filsystemet ett steg längre genom att även koppla bort lagring och hantering av filsystemets metadata. HopsFS är en öppen källkod, nästa generations distribution av Apache Hadoop Distribuerade Filsystem (HDFS) som ersätter den huvudsakliga skalbarhetsflaskhalsen i HDFS, en nod som lagrar all metadata i minnet, med en distribuerad metadatatjänst byggd på en NewSQL-databas (NDB). HopsFS lagrar filsystemets metadata fullt normaliserat i NDB, och använder sedan låsande primitiver och applikations- definierade lås för att säkerställa starkt konsistent metadata. I denna avhandling använder vi den konsistenta distribuerade hierarkiska filsys- temmetadata som tillhandahålls av HopsFS för att effektivt bygga nya klasser av applikationer som är tätt kopplade till filsystemet samt för att förbättra filsystemets interna funktioner. Först introducerar vi hbr, ett nytt blockrapporteringsprotokoll för HopsFS som tar bort en skalbarhetsflaskhals som hindrade HopsFS från att skalas till tiotusentals servrar. För det andra introducerar vi HopsFS-CL, en mycket tillgänglig molnbaserad distribution av HopsFS som distribuerar filsystemet över tillgänglighetszoner i molnet samtidigt som samma filsystemsemantik bibehålls. För det tredje introducerar vi HopsFS-S3, en mycket tillgänglig molnbaserad distribution av HopsFS som använder objektlagring som en backend för block- lagringslagret i molnet samtidigt som samma filsystemsemantik bibehålls. För det fjärde introducerar vi ePipe, en databus som båda skapar en konsistent förän- dringsström för HopsFS och så levererar korrekt beställd ström med låg latens till nedströmsklienter. Det vill säga ePipe utökar HopsFS med ett CDC-API (Change- data-capture) som inte bara ger effektiva filsystemmeddelanden utan också möjlig- gör polyglot-lagring för filsystemets metadata. Med polyglot-lagring kan vi avlasta metadatafrågor till en mer lämplig sökmotor - vi använder Elasticsearch för att tillhandahålla en fritext-sökning i filsystemets namnområde för att visa denna förmåga. Slutligen introducerar vi Hopsworks, en skalbar, projektbaserad big data-plattform om stödjer flera användare och ger stöd för samarbetsutveckling och drift för team med hjälp av utökad metadata.

Place, publisher, year, edition, pages
KTH Royal Institute of Technology, 2020
Series
TRITA-EECS-AVL ; 2020:65
National Category
Computer Systems
Research subject
Information and Communication Technology
Identifiers
urn:nbn:se:kth:diva-285872 (URN)978-91-7873-702-4 (ISBN)
Public defence
2020-12-18, https://kth-se.zoom.us/webinar/register/WN_rs7eHk0eQT2KTqefxHYcZw, Sal C, Electrum, Kistagången 16, Stockholm, 09:00 (English)
Opponent
Supervisors
Note

QC 20201111

Available from: 2020-11-11 Created: 2020-11-11 Last updated: 2022-06-25Bibliographically approved
Ismail, M., Niazi, S., Sundell, M., Ronstrom, M., Haridi, S. & Dowling, J. (2020). Distributed Hierarchical File Systems strike back in the Cloud. In: 2020 IEEE 40th international conference on distributed computing systems (ICDCS): . Paper presented at 40th IEEE International Conference on Distributed Computing Systems (ICDCS), NOV 29-DEC 01, 2020, ELECTR NETWORK (pp. 820-830). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Distributed Hierarchical File Systems strike back in the Cloud
Show others...
2020 (English)In: 2020 IEEE 40th international conference on distributed computing systems (ICDCS), Institute of Electrical and Electronics Engineers (IEEE) , 2020, p. 820-830Conference paper, Published paper (Refereed)
Abstract [en]

Cloud service providers have aligned on availability zones as an important unit of failure and replication for storage systems. An availability zone (AZ) has independent power, networking, and cooling systems and consists of one or more data centers. Multiple AZs in close geographic proximity form a region that can support replicated low latency storage services that can survive the failure of one or more AZs. Recent reductions in inter-AZ latency have made synchronous replication protocols increasingly viable, instead of traditional quorum-based replication protocols. We introduce HopsFS-CL, a distributed hierarchical file system with support for high-availability (HA) across AZs, backed by AZ-aware synchronously replicated metadata and AZ-aware block replication. HopsFS-CL is a redesign of HopsFS, a version of HDFS with distributed metadata, and its design involved making replication protocols and block placement protocols AZ-aware at all layers of its stack: the metadata serving, the metadata storage, and block storage layers. In experiments on a real-world workload from Spotify, we show that HopsFS-CL, deployed in HA mode over 3 AZs, reaches 1.66 million ops/s, and has similar performance to HopsFS when deployed in a single AZ, while preserving the same semantics.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2020
Series
IEEE International Conference on Distributed Computing Systems, ISSN 1063-6927
National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-299114 (URN)10.1109/ICDCS47774.2020.00108 (DOI)000667971400075 ()2-s2.0-85101968318 (Scopus ID)
Conference
40th IEEE International Conference on Distributed Computing Systems (ICDCS), NOV 29-DEC 01, 2020, ELECTR NETWORK
Note

QC 20210803

Not duplicate with DiVA 1467134

Available from: 2021-08-03 Created: 2021-08-03 Last updated: 2022-06-25Bibliographically approved
Ismail, M., Niazi, S., Sundell, M., Ronström, M., Haridi, S. & Dowling, J. (2020). Distributed Hierarchical File Systems strike back in the Cloud. In: : . Paper presented at 40th IEEE International Conference on Distributed Computing Systems, November 29 - December 1, 2020, Singapore.
Open this publication in new window or tab >>Distributed Hierarchical File Systems strike back in the Cloud
Show others...
2020 (English)Conference paper, Published paper (Refereed)
Abstract [en]

Cloud service providers have aligned on availability zones as an important unit of failure and replication for storage systems. An availability zone (AZ) has independent power, networking, and cooling systems and consists of one or more data centers. Multiple AZs in close geographic proximity form a region that can support replicated low latency storage services that can survive the failure of one or more AZs. Recent reductions in inter-AZ latency have made synchronous replication protocols increasingly viable, instead of traditional quorum-based replication protocols. We introduce HopsFS-CL, a distributed hierarchical file system with support for high- availability (HA) across AZs, backed by AZ-aware synchronously replicated metadata and AZ-aware block replication. HopsFS-CL is a redesign of HopsFS, a version of HDFS with distributed metadata, and its design involved making replication protocols and block placement protocols AZ-aware at all layers of its stack: the metadata serving, the metadata storage, and block storage layers. In experiments on a real-world workload from Spotify, we show that HopsFS-CL, deployed in HA mode over 3 AZs, reaches 1.66 million ops/s, and has similar performance to HopsFS when deployed in a single AZ, while preserving the same semantics.

National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-280786 (URN)
Conference
40th IEEE International Conference on Distributed Computing Systems, November 29 - December 1, 2020, Singapore
Note

QC 20210120

Available from: 2020-09-14 Created: 2020-09-14 Last updated: 2022-06-25Bibliographically approved
Ismail, M., Niazi, S., Berthou, G., Ronström, M., Haridi, S. & Dowling, J. (2020). HopsFS-S3: Extending Object Stores with POSIX-like Semantics and more. In: : . Paper presented at 21st International Middleware Conference Industrial Track, The Netherlands, 2020..
Open this publication in new window or tab >>HopsFS-S3: Extending Object Stores with POSIX-like Semantics and more
Show others...
2020 (English)Conference paper, Published paper (Other academic)
Abstract [en]

Object stores have become the de-facto platform for storage in the cloud due to their scalability, high availability, and low cost. However, they provide weaker metadata semantics and lower performance compared to distributed hierarchical file systems. In this paper, we introduce HopsFS-S3, a hybrid distributed hierarchical file system backed by an object store while preserving the file sys- tem’s strong consistency semantics. We base our implementation on HopsFS, a next-generation distribution of HDFS with distributed metadata. We redesigned HopsFS’ block storage layer to transpar- ently use an object store to store the file’s blocks without sacrificing the file system’s semantics. We also introduced a new block caching service to leverage faster NVMe storage for hot blocks. In our exper- iments, we show that HopsFS-S3 outperforms EMRFS for IO-bound workloads, with up to 20% higher performance and delivers up to 3.4𝑋 the aggregated read throughput of EMRFS. Moreover, we demonstrate that metadata operations on HopsFS-S3 (such as direc- tory rename) are up to two orders of magnitude faster than EMRFS. Finally, HopsFS-S3 opens up the currently closed metadata in ob- ject stores, enabling correctly-ordered change notifications with HopsFS’ change data capture (CDC) API and customized extensions to metadata.

National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-280792 (URN)
Conference
21st International Middleware Conference Industrial Track, The Netherlands, 2020.
Note

QC 20201118

Available from: 2020-09-14 Created: 2020-09-14 Last updated: 2022-06-25Bibliographically approved
Ismail, M., Niazi, S., Berthou, G., Ronstrom, M., Haridi, S. & Dowling, J. (2020). HopsFS-S3: Extending Object Stores with POSIX-like Semantics and more (industry track). In: Proceedings of the 2020 21st international middleware conference industrial track (Middleware industry '20): . Paper presented at 21st international middleware conference industrial track (Middleware industry '20) (pp. 23-30). Association for Computing Machinery (ACM)
Open this publication in new window or tab >>HopsFS-S3: Extending Object Stores with POSIX-like Semantics and more (industry track)
Show others...
2020 (English)In: Proceedings of the 2020 21st international middleware conference industrial track (Middleware industry '20), Association for Computing Machinery (ACM) , 2020, p. 23-30Conference paper, Published paper (Refereed)
Abstract [en]

Object stores have become the de-facto platform for storage in the cloud due to their scalability, high availability, and low cost. However, they provide weaker metadata semantics and lower performance compared to distributed hierarchical file systems. In this paper, we introduce HopsFS-S3, a hybrid distributed hierarchical file system backed by an object store while preserving the file system's strong consistency semantics. We base our implementation on HopsFS, a next-generation distribution of HDFS with distributed metadata. We redesigned HopsFS' block storage layer to transparently use an object store to store the file's blocks without sacrificing the file system's semantics. We also introduced a new block caching service to leverage faster NVMe storage for hot blocks. In our experiments, we show that HopsFS-S3 outperforms EMRFS for IO-bound workloads, with up to 20% higher performance and delivers up to 3.4X the aggregated read throughput of EMRFS. Moreover, we demonstrate that metadata operations on HopsFS-S3 (such as directory rename) are up to two orders of magnitude faster than EMRFS. Finally, HopsFS-S3 opens up the currently closed metadata in object stores, enabling correctly-ordered change notifications with HopsFS' change data capture (CDC) API and customized extensions to metadata.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2020
National Category
Computer Systems Computer Sciences
Identifiers
urn:nbn:se:kth:diva-300228 (URN)10.1145/3429357.3430521 (DOI)000684178900004 ()2-s2.0-85100505080 (Scopus ID)
Conference
21st international middleware conference industrial track (Middleware industry '20)
Note

QC 20210830

Available from: 2021-08-30 Created: 2021-08-30 Last updated: 2022-06-25Bibliographically approved
Ismail, M., Ronström, M., Haridi, S. & Dowling, J. (2019). ePipe: Near Real-Time Polyglot Persistence of HopsFS Metadata. In: 2019 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID): . Paper presented at 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, CCGRID 2019, Larnaca, Cyprus, May 14 - May 17, 2019 (pp. 92-101).
Open this publication in new window or tab >>ePipe: Near Real-Time Polyglot Persistence of HopsFS Metadata
2019 (English)In: 2019 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID), 2019, p. 92-101Conference paper, Published paper (Refereed)
Abstract [en]

Distributed OLTP databases are now used to manage metadata for distributed file systems, but they cannot also efficiently support complex queries or aggregations. To solve this problem, we introduce ePipe, a databus that both creates a consistent change stream for a distributed, hierarchical file system (HopsFS) and eventually delivers the correctly ordered stream with low latency to downstream clients. ePipe can be used to provide polyglot storage for file system metadata, allowing metadata queries to be handled by the most efficient engine for that query. For file system notifications, we show that ePipe achieves up to 56X throughput improvement over HDFS INotify and Trumpet with up to 3 orders of magnitude lower latency. For Spotify’s Hadoop workload, we show that ePipe can replicate all file system changes from HopsFS to Elasticsearch with an average replication lag of only 330 ms.

National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-254207 (URN)10.1109/CCGRID.2019.00020 (DOI)000483058700011 ()2-s2.0-85069468371 (Scopus ID)
Conference
19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, CCGRID 2019, Larnaca, Cyprus, May 14 - May 17, 2019
Note

QC 20190826

Available from: 2019-06-24 Created: 2019-06-24 Last updated: 2024-03-18Bibliographically approved
Niazi, S., Ismail, M., Haridi, S. & Dowling, J. (2019). HopsFS: Scaling Hierarchical File System Metadata Using NewSQL Databases. In: Sherif Sakr, Albert Y. Zomaya (Ed.), Encyclopedia of Big Data Technologies: (pp. 16-32). Springer
Open this publication in new window or tab >>HopsFS: Scaling Hierarchical File System Metadata Using NewSQL Databases
2019 (English)In: Encyclopedia of Big Data Technologies / [ed] Sherif Sakr, Albert Y. Zomaya, Springer, 2019, p. 16-32Chapter in book (Refereed)
Abstract [en]

Modern NewSQL database systems can be used to store fully normalized metadata for distributed hierarchical file systems, and provide high throughput and low operational latencies for the file system operations.

Place, publisher, year, edition, pages
Springer, 2019
National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-254240 (URN)10.1007/978-3-319-77525-8_146 (DOI)2-s2.0-105009177322 (Scopus ID)
Note

Part of ISBN 978-3-319-77524-1, 978-3-319-77525-8

QC 20190825

Available from: 2019-06-24 Created: 2019-06-24 Last updated: 2025-07-10Bibliographically approved
Ismail, M., Bonds, A., Niazi, S., Haridi, S. & Dowling, J. (2019). Scalable Block Reporting for HopsFS. In: 2019 IEEE International Congress on Big Data (BigData Congress): . Paper presented at IEEE International Congress on Big Data, IEEE BigData Congress 2019, Milan, Italy, July 8- July 13, 2019 (pp. 157-164).
Open this publication in new window or tab >>Scalable Block Reporting for HopsFS
Show others...
2019 (English)In: 2019 IEEE International Congress on Big Data (BigData Congress), 2019, p. 157-164Conference paper, Published paper (Refereed)
Abstract [en]

Distributed hierarchical file systems typically de- couple the storage of the file system’s metadata from the data (file system blocks) to enable the scalability of the file system. This decoupling, however, requires the introduction of a periodic synchronization protocol to ensure the consistency of the file system’s metadata and its blocks. Apache HDFS and HopsFS implement a protocol, called block reporting, where each data server periodically sends ground truth information about all its file system blocks to the metadata servers, allowing the metadata to be synchronized with the actual state of the data blocks in the file system. The network and processing overhead of the existing block reporting protocol, however, increases with cluster size, ultimately limiting cluster scalability. In this paper, we introduce a new block reporting protocol for HopsFS that reduces the protocol bandwidth and processing overhead by up to three orders of magnitude, compared to HDFS/HopsFS’ existing protocol. Our new protocol removes a major bottleneck that prevented HopsFS clusters scaling to tens of thousands of servers.

National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-254924 (URN)10.1109/BigDataCongress.2019.00035 (DOI)000539493700022 ()2-s2.0-85072782170 (Scopus ID)
Conference
IEEE International Congress on Big Data, IEEE BigData Congress 2019, Milan, Italy, July 8- July 13, 2019
Note

QC 20190902

Part of ISBN 978-1-7281-2771-2

Available from: 2019-07-09 Created: 2019-07-09 Last updated: 2024-10-22Bibliographically approved
Niazi, S., Ismail, M., Haridi, S., Dowling, J., Grohsschmiedt, S. & Ronström, M. (2017). HopsFS: Scaling Hierarchical File System Metadata Using NewSQL Databases. In: 15th USENIX Conference on File and Storage Technologies, FAST 2017, Santa Clara, CA, USA, February 27 - March 2, 2017: . Paper presented at 15th USENIX Conference on File and Storage Technologies, FAST 2017, Santa Clara, CA, USA, February 27 - March 2, 2017 (pp. 89-103). USENIX Association
Open this publication in new window or tab >>HopsFS: Scaling Hierarchical File System Metadata Using NewSQL Databases
Show others...
2017 (English)In: 15th USENIX Conference on File and Storage Technologies, FAST 2017, Santa Clara, CA, USA, February 27 - March 2, 2017, USENIX Association , 2017, p. 89-103Conference paper, Published paper (Refereed)
Abstract [en]

Recent improvements in both the performance and scalability of shared-nothing, transactional, in-memory NewSQL databases have reopened the research question of whether distributed metadata for hierarchical file systems can be managed using commodity databases. In this paper, we introduce HopsFS, a next generation distribution of the Hadoop Distributed File System (HDFS) that replaces HDFS’ single node in-memory metadata service, with a distributed metadata service built on a NewSQL database. By removing the metadata bottleneck, HopsFS enables an order of magnitude larger and higher throughput clusters compared to HDFS. Metadata capacity has been increased to at least 37 times HDFS’ capacity, and in experiments based on a workload trace from Spotify, we show that HopsFS supports 16 to 37 times the throughput of Apache HDFS. HopsFS also has lower latency for many concurrent clients, and no downtime during failover. Finally, as metadata is now stored in a commodity database, it can be safely extended and easily exported to external systems for online analysis and free-text search.

Place, publisher, year, edition, pages
USENIX Association, 2017
National Category
Engineering and Technology
Identifiers
urn:nbn:se:kth:diva-205355 (URN)000427295900007 ()2-s2.0-85077195449 (Scopus ID)
Conference
15th USENIX Conference on File and Storage Technologies, FAST 2017, Santa Clara, CA, USA, February 27 - March 2, 2017
Funder
EU, FP7, Seventh Framework Programme, 317871Swedish Foundation for Strategic Research , E2E-Clouds
Note

QC 20170424

Available from: 2017-04-13 Created: 2017-04-13 Last updated: 2022-06-27Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-6578-3902

Search in DiVA

Show all publications