PROCAT: Product Catalogue Dataset for Implicit Clustering, Permutation Learning and Structure Prediction
- Mateusz Jurewicz,
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewHost publication Subtitle
Datasets and Benchmarks TrackOriginal language
EnglishPublication milestones
- Published - 01/12/2021
Publication status
Published - 01/12/2021
Edition
2021Volume
1Host publication title
Thirty-fifth Conference on Neural Information Processing SystemsAbstract
In this dataset paper we introduce PROCAT, a novel e-commerce dataset containing expertly designed product catalogues consisting of individual product offers grouped into complementary sections. We aim to address the scarcity of existing datasets in the area of set-to-sequence machine learning tasks, which involve complex structure prediction. The task's difficulty is further compounded by the need to place into sequences rare and previously-unseen instances, as well as by variable sequence lengths and substructures, in the form of diversely-structured catalogues. PROCAT provides catalogue data consisting of over 1.5 million set items across a 4-year period, in both raw text form and with pre-processed features containing information about relative visual placement. In addition to this ready-to-use dataset, we include baseline experimental results on a proposed benchmark task from a number of joint set encoding and permutation learning model architectures.
Access to documents
Related Event
Title
Conference on Neural Information Processing Systems
Event type
ConferenceLinks
Degree of recognition
International eventDate
06/12/2021 - 14/12/2021Location
VirtualVIRTUAL
