When Bigger Is Not Better: LLM Scale and User Perception in Short-Duration Social Human-Robot Interaction
- ,
- Amanda Rasille Røn Volf
Research Output:
Working paper
Preprint
Open access
Publication Information
Output type
Research Output:
Working paper
Preprint
Original language
EnglishPublication milestones
- Published - 08/2026
Publication status
Published - 08/2026
Publisher
IEEE, United StatesBook series
- Book series name: Proceedings of the WRC Symposium on Advanced Robotics and Automation (WRC SARA)
ISSN: 2835-3358
Abstract
As Large Language Models increasingly power embodied agents, it remains unclear whether computational scaling directly enhances user perception during brief encounters. This paper investigates the impact of model parameter size 4B, 8B, and 30B on perceived intelligence and likability during short-duration, open-ended interactions. These exchanges approximate the brief social encounters found in many service-robot applications. Through a within-subjects study (N=19), we found no statistically significant overall preference for the 30B model over the 4B or 8B variants in intelligence, naturalness, enjoyment, or humor. A significant negative correlation was observed between AI interaction frequency and intelligence ranking for the 30B model (p=.005), suggesting that more experienced users may be more sensitive to differences in model capability. Within the tested interaction regime, the results indicate that increasing parameter scale alone did not produce a detectable perceptual advantage. Interaction-level dynamics such as conversational flow, responsiveness, and socially appropriate behavior may therefore be at least as important as model scale during brief, low-stakes social encounters.
