Calibrating Large Language Models Using Their Generations Only
- Dennis Thomas Ulmer,
- Martin Gubri,
- Hwaran Lee,
- Sangdoo Yun,
- Seong Joon Oh
- ,
- Parameter Lab,
- Naver AI Lab,
- Eberhard Karls University of Tübingen,
- Tübingen AI Center
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 15440–15459Publication milestones
- Published - 08/2024
Publication status
Published - 08/2024
Place of publication
BangkokVolume
Volume 1: Long PapersPublisher
Association for Computational Linguistics, United StatesPublication IDs
- Scopus: 85204485803
Host publication title
Proceedings of the 62nd Annual Meeting of the Association for Computational LinguisticsHost publication editors
- Lun-Wei Ku
- Andre Martins
- Vivek Srikumar
Abstract
As large language models (LLMs) are increasingly deployed in user-facing applications, building trust and maintaining safety by accurately quantifying a model’s confidence in its prediction becomes even more important. However, finding effective ways to calibrate LLMs—especially when the only interface to the models is their generated text—remains a challenge. We propose APRICOT (Auxiliary prediction of confidence targets): A method to set confidence targets and train an additional model that predicts an LLM’s confidence based on its textual input and output alone. This approach has several advantages: It is conceptually simple, does not require access to the target model beyond its output, does not interfere with the language generation, and has a multitude of potential usages, for instance by verbalizing the predicted confidence or using it to re-prompting the LLM to accurately reflecting its uncertainty. We show how our approach performs competitively in terms of calibration error for white-box and black-box LLMs on closed-book question-answering to detect incorrect LLM answers.
Publication metrics
PlumX, opens in new tab
Citations
28
Captures
23
Access to documents
Final published version
License:CC BY, opens in new tab
Related Event
Title
Conference on Association for Computational Linguistics
Event type
ConferenceDegree of recognition
International eventDate
11/08/2024 - 16/08/2024Location
BangkokThailand
