Skip to search boxSkip to navigationSkip to main content

IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages.

  • Aman Kumar
    ,
  • Himani Shrotriya
    ,
  • Prachi Sahu
    ,
  • Amogh Mishra
    ,
  • Raj Dabre
    ,
  • Indian Institute of Technology Madras
    ,
  • Columbia University
    ,
  • National Institute Of Information And Communications Technology, Japan
    ,
  • University of Edinburgh
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 5363-5394

Publication milestones

  • Published - 2022

Publication status

Published - 2022

Publisher

Association for Computational Linguistics, United States

Publication IDs

  • Scopus: 85144990949

Host publication title

Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing

Abstract

Natural Language Generation (NLG) for non-English languages is hampered by the scarcity of datasets in these languages. We present the IndicNLG Benchmark, a collection of datasets for benchmarking NLG for 11 Indic languages. We focus on five diverse tasks, namely, biography generation using Wikipedia infoboxes, news headline generation, sentence summarization, paraphrase generation and, question generation. We describe the created datasets and use them to benchmark the performance of several monolingual and multilingual baselines that leverage pre-trained sequence-to-sequence models. Our results exhibit the strong performance of multilingual language-specific pre-trained models, and the utility of models trained on our dataset for other related NLG tasks. Our dataset creation methods can be easily applied to modest-resource languages as they involve simple steps such as scraping news articles and Wikipedia infoboxes, light cleaning, and pivoting through machine translation data. To the best of our knowledge, the IndicNLG Benchmark is the first NLG benchmark for Indic languages and the most diverse multilingual NLG dataset, with approximately 8M examples across 5 tasks and 11 languages. The datasets and models will be publicly available.

Publication metrics

PlumX, opens in new tab

Citations
34
Captures
41

Related Event

Title

Conference on Empirical Methods in Natural Language Processing

Event type

Conference

Degree of recognition

International event

Date

07/12/2022 - 11/12/2022

Location

Abu DhabiUnited Arab Emirates