Skip to search boxSkip to navigationSkip to main content

PROCAT: Product Catalogue Dataset for Implicit Clustering, Permutation Learning and Structure Prediction

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Host publication Subtitle

Datasets and Benchmarks Track

Original language

English

Publication milestones

  • Published - 01/12/2021

Publication status

Published - 01/12/2021

Edition

2021

Volume

1

Host publication title

Thirty-fifth Conference on Neural Information Processing Systems

Abstract

In this dataset paper we introduce PROCAT, a novel e-commerce dataset containing expertly designed product catalogues consisting of individual product offers grouped into complementary sections. We aim to address the scarcity of existing datasets in the area of set-to-sequence machine learning tasks, which involve complex structure prediction. The task's difficulty is further compounded by the need to place into sequences rare and previously-unseen instances, as well as by variable sequence lengths and substructures, in the form of diversely-structured catalogues. PROCAT provides catalogue data consisting of over 1.5 million set items across a 4-year period, in both raw text form and with pre-processed features containing information about relative visual placement. In addition to this ready-to-use dataset, we include baseline experimental results on a proposed benchmark task from a number of joint set encoding and permutation learning model architectures.

Related Event

Title

Conference on Neural Information Processing Systems

Event type

Conference

Degree of recognition

International event

Date

06/12/2021 - 14/12/2021

Location

VirtualVIRTUAL