Skip to search boxSkip to navigationSkip to main content

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 1-17 (17 pages)

Publication milestones

  • Published - 13/04/2026

Publication status

Published - 13/04/2026

Place of publication

New York, NY, USA

Publisher

Association for Computing Machinery, United States
9798400722783

ISBN (Electronic)

979-8-4007-2278-3

Publication IDs

  • Scopus: 105038655976

Host publication title

Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems

Abstract

How do product teams evaluate LLM-powered products? As organizations integrate large language models (LLMs) into digital products, their unpredictable nature makes traditional evaluation approaches inadequate, yet little is known about how practitioners navigate this challenge. Through interviews with nineteen practitioners across diverse sectors, we identify ten evaluation practices spanning informal 'vibe checks' to organizational meta-work. Beyond confirming four documented challenges, we introduce a novel fifth we call the results-actionability gap, in which practitioners gather evaluation data but cannot translate findings into concrete improvements. Drawing on patterns from successful teams, we contribute strategies to bridge this gap, supporting practitioners' formalization journey from ad-hoc interpretive practices (e.g., vibe checks) toward systematic evaluation. Our analysis suggests these interpretive practices are necessary adaptations to LLM characteristics rather than methodological failures. For HCI researchers, this presents a research opportunity to support practitioners in systematizing emerging practices rather than developing new evaluation frameworks.

Publication metrics

Funding Details

This research was partially funded by Danish Novo Nordisk Foundation under Grant Number NNF20OC0066119 and the Science of trustworthy AI award from Schmidt Sciences.

Related Event

Title

Conference on Human Factors in Computing Systems

Event type

Conference

Degree of recognition

International event

Date

13/04/2026 - 17/04/2026

Location

Centre de Convencions Internacional de Barcelona.BarcelonaSpain