EARBench: Towards Evaluating Physical Risk Awareness for Task Planning of Foundation Model-based Embodied AI Agents

Baoyuan Wu
Zihao Zhu
Bingzhe Wu
Zhengyou ZHANG
Lei Han
Qingshan Liu

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Embodied artificial intelligence (EAI) integrates advanced Al models into physical entities for real-world interaction. The emergence of foundation models as the "brain" of EAI agents for high-level task planning has shown promising results. However, the deployment of these agents in physical environments presents significant safety challenges. For instance, a housekeeping robot lacking sufficient risk awareness might place a metal container in a microwave, potentially causing a fire. To address these critical safety concerns, comprehensive pre-deployment risk assessments are imperative. This study introduces the first embodied AI risk benchmark (EARBench), a novel framework for automated physical risk assessment in EAI scenarios. EARBench employs a multi-agent cooperative system that leverages various foundation models to generate safety guidelines, create risk-prone scenarios, make task planning, and evaluate safety systematically. Utilizing this framework, we construct EARDataset, comprising diverse test cases across various domains, encompassing both textual and visual scenarios. Our comprehensive evaluation of state-of-the-art foundation models reveals alarming results: all models exhibit high task risk rates (TRR), with an average of 95.75% across all evaluated models. Notably, even GPT-4o, widely regarded as one of the most advanced models, demonstrates a TRR of 94.03%, underscoring the pervasive lack of risk identification and avoidance capabilities in complex physical environments. To address these challenges, we further propose two prompting-based risk mitigation strategies. While these strategies demonstrate some efficacy in reducing TRR, the improvements are limited, still indicating substantial safety concerns. This study provides the first large-scale assessment of physical risk awareness in EAI agents. Our findings underscore the critical need for enhanced safety measures in EAI systems and provide valuable insights for future research directions in developing safer embodied artificial intelligence system. Data and code are available at https://github.com/zihao-ai/EARBench.

Version published to 10.21203/rs.3.rs-5540665/v1 on Research Square
Feb 3, 2025

EAISE: A Simulation Environment for Self-Evolving Embodied AI with Mirror Testing and Multi-Agent Diagnostics

This article has 1 author:
1. Berend Watchus
This article has no evaluationsLatest version Jun 20, 2025
Safe Reinforcement Learning for Vision-Based Robotic Manipulation in Human-Centered Environments

This article has 7 authors:
1. Fawad Khan
2. Wei Feng
3. Zhiyong Wang
4. Tianlun Huang
5. Xiao Liu
6. Yunduan Cui
7. Wang Weijun
This article has no evaluationsLatest version Jun 16, 2025
Advancing Conversational Diagnostic AI with Multimodal Reasoning

This article has 36 authors:
1. Ryutaro Tanno
2. Khaled Saab
3. Jan Freyberg
4. Chunjong Park
5. Tim Strother
6. Yong Cheng
7. Wei-Hung Weng
8. David Barrett
9. David Stutz
10. Nenad Tomasev
11. Anil Palepu
12. Valentin Liévin
13. Yash Sharma
14. Abdullah Ahmed
15. Elahe Vedadi
16. Roma Ruparel
17. Kimberly Kanada
18. Cian Hughes
19. Yun Liu
20. Geoff Brown
21. Yang Gao
22. Sean Li
23. S. Sara Mahdavi
24. James Manyika
25. Katherine Chou
26. Yossi Matias
27. Avinatan Hassidim
28. Dale Webster
29. Pushmeet Kohli
30. Ali Eslami
31. Joelle Barral
32. Adam Rodman
33. Vivek Natarajan
34. Mike Schaekermann
35. Tao Tu
36. Alan Karthikesalingam
This article has no evaluationsLatest version Jun 4, 2025

Listed in

Abstract

Article activity feed

Related articles

EAISE: A Simulation Environment for Self-Evolving Embodied AI with Mirror Testing and Multi-Agent Diagnostics

Safe Reinforcement Learning for Vision-Based Robotic Manipulation in Human-Centered Environments

Advancing Conversational Diagnostic AI with Multimodal Reasoning