Skip to search boxSkip to navigationSkip to main content

Towards Engineering a Web-Scale Multimedia Service: A Case Study Using Spark

  • Reykjavík University
    ,
  • Research Institute Computer And Systems Aléatoires
    ,
  • The University of Chicago
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 1-12

Publication milestones

  • Published - 06/2017

Publication status

Published - 06/2017

Place of publication

Taipei, Taiwan

Publisher

Association for Computing Machinery, United States
978-1-4503-5002-0

Publication IDs

  • Scopus: 85025671251

Host publication title

Proceedings of the ACM Multimedia Systems Conference (MMSys)

Abstract

Computing power has now become abundant with multi-core machines, grids and clouds, but it remains a challenge to harness the available power and move towards gracefully handling web-scale datasets. Several researchers have used automatically distributed computing frameworks, notably Hadoop and Spark, for processing multimedia material, but mostly using small collections on small clusters. In this paper, we describe the engineering process for a prototype of a (near) web-scale multimedia service using the Spark framework running on the AWS cloud service. We present experimental results using up to 43 billion SIFT feature vectors from the public YFCC 100M collection, making this the largest high-dimensional feature vector collection reported in the literature. The design of the prototype and performance results demonstrate both the flexibility and scalability of the Spark framework for implementing multimedia services.

Publication metrics

PlumX

Captures
11
Citations
14

Access to documents