spot_img
HomeResearch & DevelopmentDetecting Robot Errors Through Human Signals: The ERR@HRI 2.0...

Detecting Robot Errors Through Human Signals: The ERR@HRI 2.0 Challenge

TLDR: The ERR@HRI 2.0 Challenge provides a comprehensive multimodal dataset and framework for researchers to develop machine learning models that detect errors and failures in human-robot conversations. It focuses on identifying robot errors from both the robot’s system perspective and the user’s perspective, leveraging facial, speech, and head movement data collected from 16 hours of interactions with LLM-powered conversational robots. The challenge aims to advance the field of human-robot interaction by improving real-time error detection through social signal analysis.

As large language models (LLMs) become increasingly integrated into conversational robots, our interactions with these machines are becoming more dynamic and natural. However, even with advanced LLMs, these robots are still prone to making mistakes. These errors can range from misunderstanding what a user intends, interrupting conversations prematurely, or sometimes failing to respond at all. Detecting and addressing these failures is crucial to prevent conversations from breaking down, avoid disrupting tasks, and maintain user trust in the robot.

To tackle this significant challenge, the ERR@HRI 2.0 Challenge was launched. This initiative provides a unique multimodal dataset specifically designed to help researchers develop and benchmark machine learning models capable of detecting robot failures during human-robot conversations. The dataset is extensive, comprising 16 hours of two-way human-robot interactions. It captures a rich array of data, including facial expressions, speech patterns, and head movements from the human participants. Each interaction is carefully annotated, indicating the presence or absence of robot errors from the robot’s own system perspective, as well as the user’s perceived intention to correct a mismatch between the robot’s behavior and their expectations.

Understanding Robot Errors from Two Angles

The challenge uniquely defines robot errors from two distinct perspectives: the system perspective and the user perspective. A robot error from the system perspective refers to a noticeable deviation in the robot’s behavior from its intended design. This could include the robot failing to understand user intent, interrupting the user, or providing an inappropriate response. On the other hand, a robot error from the user perspective is defined by observable verbal and non-verbal user-initiated disruptive interruptions. These are actions taken by the user to signal an intention to correct a mismatch between the robot’s behavior and what they expected.

The dataset used in the challenge includes interactions with two types of LLM-powered robots: a social robot with more human-like features and expressions, and a smart speaker. Participants engaged with these robots in various tasks, such as medical self-diagnosis, trip planning, a desert survival simulation, and discussions on topics like capital punishment or university police forces. To ensure user privacy, the challenge provides extracted multimodal behavioral features rather than raw audio and video recordings. These features include detailed facial and head pose data (from OpenFace 2.2.0), comprehensive audio features (from openSMILE toolbox), and transcribed speech features, including speaker diarization and CLIP embeddings.

Also Read:

Evaluating Performance and Future Steps

Participants in the ERR@HRI 2.0 Challenge are invited to form teams and develop machine learning models using this rich multimodal data. Submissions are evaluated using various performance metrics, including detection accuracy and false positive rates, with a strong emphasis on the overall F1 score for on-the-fly evaluation, which simulates real-time performance. The dataset is split into training and test sets using a subject-independent strategy to prevent models from overfitting. Baseline models, using Random Forest and techniques like SMOTE to handle data imbalance, were provided to guide participants.

This challenge represents a significant stride toward improving failure detection in human-robot interaction through the analysis of social signals. By bringing together researchers from multimedia, robotics, and human-robot interaction communities, the ERR@HRI 2.0 Challenge aims to advance our understanding of how multimodal data can enhance the analysis and interpretation of interactions between humans and autonomous robots. For more detailed information, you can refer to the full research paper here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -