[Ph.D Thesis] Active Clarification in Human-Robot Interaction via Object State Perception and Reasoning
Published in Staats- und Universitätsbibliothek Hamburg Carl von Ossietzky (SUB Hamburg), 2026
Abstract: Embodied agents have benefited from advances in Artificial Intelligence (AI), enabling them to operate in human-centered environments such as domestic or restaurant settings. In such environments, Human-Robot Interaction (HRI) is characterized by Uncertainty, Vagueness, and Ambiguity, particularly when interaction refers to fine-grained object states (e.g., clean vs. dirty) rather than object categories alone. This increased granularity introduces significant challenges for embodied agents in integrating language, perception, and action.
This thesis addresses these challenges by investigating how embodied agents can acquire and process object states for robust communication and interaction. It begins with a comprehensive review of existing approaches from an architectural perspective, synthesizing recent trends that have advanced the development of communication models, conversational agents, and transformer-based language and vision-language models. Building on this foundation, the thesis first clarifies the concepts of uncertainty, vagueness, and ambiguity within the interdisciplinary domains of linguistics and HRI, and then discusses object states. This analysis motivates four methodological contributions, each addressing different combinations of the identified conceptual, representational, and methodological problems.
First, a novel HYbrid Neural Architecture (HYNA) is introduced to advance visually grounded human–robot dialog by resolving ambiguities between visual context and natural language. Next, the Object State-Sensitive Agent (OSSA) is proposed to evaluate and improve the performance of large language models (LLMs) and vision-language models (VLMs) in robotic task planning. OSSA incorporates object state representations and commonsense knowledge to support long-horizon instruction execution in dynamic environments, including user-adaptive and multimodal interaction scenarios. Additionally, the Dialog Bridge mechanism is formulated to assess LLMs’ capabilities in user age-sensitive adaptation and multimodal processing in HRI. Finally, the State-sensitive Vision-Language Model (StateVLM) is presented to enhance the numerical reasoning capabilities of VLMs. StateVLM is improved through an efficient fine-tuning strategy with an auxiliary regression loss, enabling better object localization and affordance reasoning at the object state-level with limited annotated data and computational resources.
Overall, this thesis advances embodied intelligence by integrating knowledge-driven reasoning with visual perception and language models, thereby improving communication and visual understanding for embodied agents operating in dynamic real-world HRI scenarios.
Recommended citation: Sun, X. W. “Active clarification in human-robot interaction via object state perception and reasoning,” Ph.D. dissertation, Universität Hamburg, Hamburg, Germany, 2026. [Online]. Available: https://ediss.sub.uni-hamburg.de/handle/ediss/12538.
Download Paper
