Skip to search boxSkip to navigationSkip to main content

From Hearing to Understanding: Uncertainty-Aware Social Context Inference for Agentic Robots

Research Output:
Working paper
Preprint

Open access

Publication Information

Output type

Research Output:
Working paper
Preprint

Original language

English

Publication milestones

  • Published - 08/2026

Publication status

Published - 08/2026

Publisher

IEEE, United States

Book series

  • Book series name: Proceedings of the International Conference on Robotics and Automation (ICRA)

Abstract

For robots to act appropriately in human spaces, they must read the social context they have entered, not merely perceive its objects. We present an agentic system that infers social context as an incremental, uncertainty-aware belief over six interpretable dimensions (collaboration, competition, empathy, hierarchy, formality, and emotional tone), refining that belief across three stages: prosodic and transcribed audio, a scene image, and a clarifying question posed to a bystander. Motivated by how people form an expectation of a situation from what they hear before they see it, the pipeline primes its visual interpretation with pre-experienced audio. We evaluated the system in a within-subjects study against an aggregated human baseline (N = 22) on four social situations, comparing a locally hosted model configuration against a higher-capacity commercial one. Audio alone recovered most of the human rating pattern and predicted the final human visual consensus as well as any later stage (r = .82); neither visual refinement nor the clarifying question yielded a reliable aggregate gain in accuracy. Most notably, the system's reported uncertainty collapsed by 52% across stages while its error did not fall and human uncertainty barely changed, leaving the reported confidence three to seven times narrower than the actual error; greater model capacity lowered error but not this overconfidence. These results indicate that pre-visual audio priming is a powerful, low-cost signal for social-context inference, and that the central open challenge is calibration rather than raw accuracy.