Skip to search boxSkip to navigationSkip to main content

Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics

  • University of Texas
    ,
  • Tokyo Institute of Technology
    ,
  • RIKEN Center for Computational Science
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 125-135 (11 pages)

Publication milestones

  • Published - 01/11/2021

Publication status

Published - 01/11/2021

Place of publication

Online and Punta Cana, Dominican Republic

Publisher

Association for Computational Linguistics, United States

Host publication title

Proceedings of the Second Workshop on Insights from Negative Results in NLP

Abstract

Much of recent progress in NLU was shown to be due to models' learning dataset-specific heuristics. We conduct a case study of generalization in NLI (from MNLI to the adversarially constructed HANS dataset) in a range of BERT-based architectures (adapters, Siamese Transformers, HEX debiasing), as well as with subsampling the data and increasing the model size. We report 2 successful and 3 unsuccessful strategies, all providing insights into how Transformer-based models learn to generalize.

Related Event

Title

Insights from Negative Results in NLP

Event type

Conference

Date

01/11/2021

Location

VIRTUALDominican Republic