Skip to search boxSkip to navigationSkip to main content

When BERT Plays the Lottery, All Tickets Are Winning

  • Zoho
    ,
  • University of Massachusetts
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 3208-3229 (22 pages)

Publication milestones

  • Published - 01/11/2020

Publication status

Published - 01/11/2020

Place of publication

Online

Publisher

Association for Computational Linguistics, United States

Publication IDs

  • Scopus: 85102547740

Host publication title

Proceedings of EMNLP

Abstract

Much of the recent success in NLP is due to the large Transformer-based models such as BERT (Devlin et al, 2019). However, these models have been shown to be reducible to a smaller number of self-attention heads and layers. We consider this phenomenon from the perspective of the lottery ticket hypothesis. For fine-tuned BERT, we show that (a) it is possible to find a subnetwork of elements that achieves performance comparable with that of the full model, and (b) similarly-sized subnetworks sampled from the rest of the model perform worse. However, the "bad" subnetworks can be fine-tuned separately to achieve only slightly worse performance than the "good" ones, indicating that most weights in the pre-trained BERT are potentially useful. We also show that the "good" subnetworks vary considerably across GLUE tasks, opening up the possibilities to learn what knowledge BERT actually uses at inference time.

Publication metrics

PlumX

Citations
119
Captures
246

Related Event

Title

Conference on Empirical Methods in Natural Language Processing

Event type

Conference

Date

16/11/2020 - 20/11/2020

Location

OnlineVIRTUAL