Data Pipes: Declarative Control over Data Movement
- Lukas Vogel,
- Daniel Ritter,
- Danica Porobic,
- ,
- Tianzheng Wang,
- Alberto Lerner
- Technical University of Munich,
- SAP Research,
- Oracle Corporation,
- ,
- ,
- Simon Fraser University
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPublication milestones
- Published - 2023
Publication status
Published - 2023
Host publication title
Conference on Innovative Data Systems ResearchAbstract
Today’s storage landscape offers a deep and heterogeneous stack of technologies that promises to meet even the most demanding data-intensive workload needs. The diversity of technologies, however, presents a challenge. Parts of it are not controlled directly by the application, e.g., the cache layers, and the parts that are controlled, often require the programmer to deal with very different transfer mechanisms, such as disk and network APIs. Combining these
different abstractions properly require great skill, and even so, expert-written programs can lead to sub-optimal utilization of the storage stack and present performance unpredictability.
In this paper, we propose to combat these issues with a new programming abstraction called Data Pipes. Data pipes offer a new API that can express data transfers uniformly, irrespective of the source and destination data placements. By doing so, they can orchestrate how data moves over the different layers of the storage stack explicitly and fluidly. We suggest a preliminary implementation of Data Pipes that relies mainly on existing hardware primitives to implement data movements. We evaluate this implementation experimentally and comment on how a full version of Data Pipes could be brought to fruition.
different abstractions properly require great skill, and even so, expert-written programs can lead to sub-optimal utilization of the storage stack and present performance unpredictability.
In this paper, we propose to combat these issues with a new programming abstraction called Data Pipes. Data pipes offer a new API that can express data transfers uniformly, irrespective of the source and destination data placements. By doing so, they can orchestrate how data moves over the different layers of the storage stack explicitly and fluidly. We suggest a preliminary implementation of Data Pipes that relies mainly on existing hardware primitives to implement data movements. We evaluate this implementation experimentally and comment on how a full version of Data Pipes could be brought to fruition.
Access to documents
Final published version
Related Event
Title
Conference on Innovative Data Systems Research
Event type
ConferenceDegree of recognition
International eventDate
08/01/2023 - 11/01/2023Location
AmsterdamNetherlands
