What's in Your Embedding, And How It Predicts Task Performance
- ,
- Shashwath Hosur Ananthakrishna,
- Anna Rumshisky
- University of Massachusetts
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewPublication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 2690-2703 (14 pages)Publication milestones
- Published - 2018
Publication status
Published - 2018
Place of publication
Santa Fe, New Mexico, USA, August 20-26, 2018Publisher
Association for Computational Linguistics, United StatesHost publication title
Proceedings of the 27th International Conference on Computational LinguisticsAbstract
Attempts to find a single technique for general-purpose intrinsic evaluation of word embeddings have so far not been successful. We present a new approach based on scaled-up qualitative analysis of word vector neighborhoods that quantifies interpretable characteristics of a given model (e.g. its preference for synonyms or shared morphological forms as nearest neighbors). We analyze 21 such factors and show how they correlate with performance on 14 extrinsic and intrinsic task datasets (and also explain the lack of correlation between some of them). Our approach enables multi-faceted evaluation, parameter search, and generally – a more principled, hypothesis-driven approach to development of distributional semantic representations.
