Skip to search boxSkip to navigationSkip to main content

Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat

  • Ravi Shekhar(Author)
    ,
  • Aashish Venkatesh(Author)
    ,
  • Tim Baumgärtner(Author)
    ,
  • Elia Bruni(Author)
    ,
  • ,
  • Raffaella Bernardi
  • University of Trento
    ,
  • University of Amsterdam
    ,
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Publication milestones

  • Published - 06/2019

Publication status

Published - 06/2019

Place of publication

Minneapolis

Publisher

Association for Computational Linguistics, United States

ISBN (Electronic)

978-1-950737-13-0

Publication IDs

  • Scopus: 85074818683

Host publication title

NAACL (North American Association for Computational Linguistics)

Abstract

We propose a grounded dialogue state encoder which addresses a foundational issue on how to integrate visual grounding with dialogue system components. As a test-bed, we focus on the GuessWhat?! game, a two-player game where the goal is to identify an object in a complex visual scene by asking a sequence of yes/no questions. Our visually-grounded encoder lever- ages synergies between guessing and asking questions, as it is trained jointly using multi- task learning. We further enrich our model via a cooperative learning regime. We show that the introduction of both the joint architecture and cooperative learning lead to accuracy improvements over the baseline system. We compare our approach to an alternative system which extends the baseline with reinforcement learning. Our in-depth analysis shows that the linguistic skills of the two models differ dramatically, despite approaching comparable performance levels. This points at the importance of analyzing the linguistic output of competing systems beyond numeric comparison solely based on task success.

Publication metrics

PlumX

Citations
40
Captures
109