Skip to main navigation Skip to search Skip to main content

Benchmarking Column Statistics for Analytical Query Pruning

  • Snowflake

Research output: Conference Article in Proceeding or Book/Report chapterArticle in proceedingsResearchpeer-review

Abstract

Min/max indexes are the most widely used statistics for partition pruning in analytical systems, but they miss pruning opportunities when partitions contain outliers or wide value ranges. Richer statistics, including multiple min/max indexes per partition, Bloom filters, and dictionaries, can improve pruning effectiveness at the cost of additional metadata, yet no study quantifies when this overhead is justified. We present a benchmark for evaluating column statistics for partition pruning, implemented in DuckDB, and assess pruning effectiveness and query runtimes across metadata budgets of 10 × to 1000 × the size of standard min/max indexes. Our results show that richer statistics yield gains for unsorted data, data with outliers, and low-cardinality columns. We derive initial guidelines for selecting column statistics based on data characteristics at load time.
Original languageEnglish
Title of host publicationDBTest '26: Proceedings of the 2026 11th International Workshop on Testing Database Systems
Number of pages31
PublisherAssociation for Computing Machinery
Publication date26 Jun 2026
Pages7
ISBN (Print) 979-8-4007-2701
DOIs
Publication statusPublished - 26 Jun 2026
Event11th International Workshop on Testing Database Systems
- Bengaluru, India
Duration: 5 Jun 20265 Jun 2026

Workshop

Workshop11th International Workshop on Testing Database Systems
Country/TerritoryIndia
CityBengaluru
Period05/06/202605/06/2026
SeriesProceedings of the International Workshop on Testing Database Systems

Keywords

  • Column statistics
  • Partition pruning
  • Data skipping

Fingerprint

Dive into the research topics of 'Benchmarking Column Statistics for Analytical Query Pruning'. Together they form a unique fingerprint.

Cite this