Skip to search boxSkip to navigationSkip to main content

When Bigger Is Not Better: LLM Scale and User Perception in Short-Duration Social Human-Robot Interaction

Research Output:
Working paper
Preprint

Open access

Publication Information

Output type

Research Output:
Working paper
Preprint

Original language

English

Publication milestones

  • Published - 08/2026

Publication status

Published - 08/2026

Publisher

IEEE, United States

Book series

  • Book series name: Proceedings of the WRC Symposium on Advanced Robotics and Automation (WRC SARA)
    ISSN: 2835-3358

Abstract

As Large Language Models increasingly power embodied agents, it remains unclear whether computational scaling directly enhances user perception during brief encounters. This paper investigates the impact of model parameter size 4B, 8B, and 30B on perceived intelligence and likability during short-duration, open-ended interactions. These exchanges approximate the brief social encounters found in many service-robot applications. Through a within-subjects study (N=19), we found no statistically significant overall preference for the 30B model over the 4B or 8B variants in intelligence, naturalness, enjoyment, or humor. A significant negative correlation was observed between AI interaction frequency and intelligence ranking for the 30B model (p=.005), suggesting that more experienced users may be more sensitive to differences in model capability. Within the tested interaction regime, the results indicate that increasing parameter scale alone did not produce a detectable perceptual advantage. Interaction-level dynamics such as conversational flow, responsiveness, and socially appropriate behavior may therefore be at least as important as model scale during brief, low-stakes social encounters.