Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
- ,
- Malak Sadek,
- Ziang Xiao,
- ,
- Q. Vera Liao,
- ,
- ,
- University of Cambridge,
- Johns Hopkins University,
- ,
- University of Michigan
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 1-17 (17 pages)Publication milestones
- Published - 13/04/2026
Publication status
Published - 13/04/2026
Place of publication
New York, NY, USAPublisher
Association for Computing Machinery, United StatesISBN (Print)
9798400722783ISBN (Electronic)
979-8-4007-2278-3Publication IDs
- Scopus: 105038655976
Host publication title
Proceedings of the 2026 CHI Conference on Human Factors in Computing SystemsAbstract
How do product teams evaluate LLM-powered products? As organizations integrate large language models (LLMs) into digital products, their unpredictable nature makes traditional evaluation approaches inadequate, yet little is known about how practitioners navigate this challenge. Through interviews with nineteen practitioners across diverse sectors, we identify ten evaluation practices spanning informal 'vibe checks' to organizational meta-work. Beyond confirming four documented challenges, we introduce a novel fifth we call the results-actionability gap, in which practitioners gather evaluation data but cannot translate findings into concrete improvements. Drawing on patterns from successful teams, we contribute strategies to bridge this gap, supporting practitioners' formalization journey from ad-hoc interpretive practices (e.g., vibe checks) toward systematic evaluation. Our analysis suggests these interpretive practices are necessary adaptations to LLM characteristics rather than methodological failures. For HCI researchers, this presents a research opportunity to support practitioners in systematizing emerging practices rather than developing new evaluation frameworks.
Publication metrics
PlumX, opens in new tab
Captures
8
Funding Details
This research was partially funded by Danish Novo Nordisk Foundation under Grant Number NNF20OC0066119 and the Science of trustworthy AI award from Schmidt Sciences.
Access to documents
Accepted author manuscript, 580.26 KB
License:CC BY, opens in new tab
Related Event
Title
Conference on Human Factors in Computing Systems
Event type
ConferenceDegree of recognition
International eventDate
13/04/2026 - 17/04/2026Location
Centre de Convencions Internacional de Barcelona.BarcelonaSpain
