Skip to search boxSkip to navigationSkip to main content

Large-Scale Similarity Joins With Guarantees

  • Rasmus Pagh
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 15-24

Publication milestones

  • Published - 2015

Publication status

Published - 2015

Volume

31

Book series

  • Book series name: Leibniz International Proceedings in Informatics
    ISSN: 1868-8969

ISBN (Electronic)

978-3-939897-79-8

Host publication title

18th International Conference on Database Theory (ICDT 2015)

Abstract

The ability to handle noisy or imprecise data is becoming increasingly important in computing. In the database community the notion of similarity join has been studied extensively, yet existing solutions have offered weak performance guarantees. Either they are based on deterministic filtering techniques that often, but not always, succeed in reducing computational costs, or they are based on randomized techniques that have improved guarantees on computational cost but come with a probability of not returning the correct result. The aim of this paper is to give an overview of randomized techniques for high-dimensional similarity search, and discuss recent advances towards making these techniques more widely applicable by eliminating probability of error and improving the locality of data access.