direkt zum Inhalt springen

direkt zum Hauptnavigationsmenü

Sie sind hier

TU Berlin

Inhalt des Dokuments

Publikationen

Muses: Distributed Data Migration System for Polystores
Zitatschlüssel DBLP:conf/icde/KaitouaRKM19
Autor Abdulrahman Kaitoua, Tilmann Rabl, Asterios Katsifodimos, Volker Markl
Jahr 2019
DOI https://doi.org/10.1109/ICDE.2019.00152
Journal 35th IEEE International Conference on Data Engineering, ICDE
Zusammenfassung Large datasets can originate from various sources and are being stored in heterogeneous formats, schemas, and locations. Typical data science tasks need to combine those datasets in order to increase their value and extract knowledge. This is done in various data processing systems with diverse execution engines. In order to take advantage of each execution engine’s characteristics and APIs data scientists need to migrate and transform their datasets at a very high computational cost and manual labor. Data migration is challenging for two main reasons: i) execution engines expect specific types/shapes of the data as input; ii) there are various physical representations of the data (e.g., partitions). Therefore, migrating data efficiently requires knowledge of systems internals and assumptions. In this paper we present Muses, a distributed, high performance data migration engine that is able to forward, transform, repartition, and broadcast data between distributed engines’ instances efficiently. Muses does not require any changes in the underlying execution engines. In an experimental evaluation, we show that migrating data from one execution engine to another (in order to take advantage of faster, native operations) can increase a pipeline’s performance by 30%. Index Terms—Distributed systems, data migration, data transformation, big data engine, data integration.
Link zur Publikation Link zur Originalpublikation Download Bibtex Eintrag

Zusatzinformationen / Extras

Direktzugang:

Schnellnavigation zur Seite über Nummerneingabe