From Hearing to Understanding: Uncertainty-Aware Social Context Inference for Agentic Robots
Research Output:
Working paper
Preprint
Open access
Publication Information
Output type
Research Output:
Working paper
Preprint
Original language
EnglishPublication milestones
- Published - 08/2026
Publication status
Published - 08/2026
Publisher
IEEE, United StatesBook series
- Book series name: Proceedings of the International Conference on Robotics and Automation (ICRA)
Abstract
For robots to act appropriately in human spaces, they must read the social context they have entered, not merely perceive its objects. We present an agentic system that infers social context as an incremental, uncertainty-aware belief over six interpretable dimensions (collaboration, competition, empathy, hierarchy, formality, and emotional tone), refining that belief across three stages: prosodic and transcribed audio, a scene image, and a clarifying question posed to a bystander. Motivated by how people form an expectation of a situation from what they hear before they see it, the pipeline primes its visual interpretation with pre-experienced audio. We evaluated the system in a within-subjects study against an aggregated human baseline (N = 22) on four social situations, comparing a locally hosted model configuration against a higher-capacity commercial one. Audio alone recovered most of the human rating pattern and predicted the final human visual consensus as well as any later stage (r = .82); neither visual refinement nor the clarifying question yielded a reliable aggregate gain in accuracy. Most notably, the system's reported uncertainty collapsed by 52% across stages while its error did not fall and human uncertainty barely changed, leaving the reported confidence three to seven times narrower than the actual error; greater model capacity lowered error but not this overconfidence. These results indicate that pre-visual audio priming is a powerful, low-cost signal for social-context inference, and that the central open challenge is calibration rather than raw accuracy.
Access to documents
Submitted manuscript, 769.28 KB
License:CC BY-NC-ND, opens in new tab
