Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics
- Prajjwal Bhargava,
- Aleksandr Drozd,
- University of Texas,
- Tokyo Institute of Technology,
- RIKEN Center for Computational Science
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewPublication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 125-135 (11 pages)Publication milestones
- Published - 01/11/2021
Publication status
Published - 01/11/2021
Place of publication
Online and Punta Cana, Dominican RepublicPublisher
Association for Computational Linguistics, United StatesHost publication title
Proceedings of the Second Workshop on Insights from Negative Results in NLPAbstract
Much of recent progress in NLU was shown to be due to models' learning dataset-specific heuristics. We conduct a case study of generalization in NLI (from MNLI to the adversarially constructed HANS dataset) in a range of BERT-based architectures (adapters, Siamese Transformers, HEX debiasing), as well as with subsampling the data and increasing the model size. We report 2 successful and 3 unsuccessful strategies, all providing insights into how Transformer-based models learn to generalize.
Related Event
Title
Insights from Negative Results in NLP
Event type
ConferenceDate
01/11/2021 Location
VIRTUALDominican Republic
