TY - UNPB
T1 - From Hearing to Understanding: Uncertainty-Aware Social Context Inference for Agentic Robots
AU - Frederiksen, Morten Roed
PY - 2026/8
Y1 - 2026/8
N2 - For robots to act appropriately in human spaces, they must read the social context they have entered, not merely perceive its objects. We present an agentic system that infers social context as an incremental, uncertainty-aware belief over six interpretable dimensions (collaboration, competition, empathy, hierarchy, formality, and emotional tone), refining that belief across three stages: prosodic and transcribed audio, a scene image, and a clarifying question posed to a bystander. Motivated by how people form an expectation of a situation from what they hear before they see it, the pipeline primes its visual interpretation with pre-experienced audio. We evaluated the system in a within-subjects study against an aggregated human baseline (N = 22) on four social situations, comparing a locally hosted model configuration against a higher-capacity commercial one. Audio alone recovered most of the human rating pattern and predicted the final human visual consensus as well as any later stage (r = .82); neither visual refinement nor the clarifying question yielded a reliable aggregate gain in accuracy. Most notably, the system's reported uncertainty collapsed by 52% across stages while its error did not fall and human uncertainty barely changed, leaving the reported confidence three to seven times narrower than the actual error; greater model capacity lowered error but not this overconfidence. These results indicate that pre-visual audio priming is a powerful, low-cost signal for social-context inference, and that the central open challenge is calibration rather than raw accuracy.
AB - For robots to act appropriately in human spaces, they must read the social context they have entered, not merely perceive its objects. We present an agentic system that infers social context as an incremental, uncertainty-aware belief over six interpretable dimensions (collaboration, competition, empathy, hierarchy, formality, and emotional tone), refining that belief across three stages: prosodic and transcribed audio, a scene image, and a clarifying question posed to a bystander. Motivated by how people form an expectation of a situation from what they hear before they see it, the pipeline primes its visual interpretation with pre-experienced audio. We evaluated the system in a within-subjects study against an aggregated human baseline (N = 22) on four social situations, comparing a locally hosted model configuration against a higher-capacity commercial one. Audio alone recovered most of the human rating pattern and predicted the final human visual consensus as well as any later stage (r = .82); neither visual refinement nor the clarifying question yielded a reliable aggregate gain in accuracy. Most notably, the system's reported uncertainty collapsed by 52% across stages while its error did not fall and human uncertainty barely changed, leaving the reported confidence three to seven times narrower than the actual error; greater model capacity lowered error but not this overconfidence. These results indicate that pre-visual audio priming is a powerful, low-cost signal for social-context inference, and that the central open challenge is calibration rather than raw accuracy.
KW - social
KW - context
KW - comprehension
KW - context aware
KW - robots
KW - context inference
KW - agentic AI
M3 - Preprint
T3 - Proceedings of the International Conference on Robotics and Automation (ICRA)
BT - From Hearing to Understanding: Uncertainty-Aware Social Context Inference for Agentic Robots
PB - IEEE
ER -