ICML GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning

Poster
in
Workshop: Workshop on Computer Use Agents

GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning

Zhen Xiang · Linzhi Zheng · Yanjie Li · Junyuan Hong · Qinbin Li · Han Xie · Jiawei Zhang · Zidi Xiong · Chulin Xie · Nathaniel Bastian · Carl Yang · Dawn Song · Bo Li

[ Abstract ] [ Project Page ]

[ OpenReview]

Abstract:

The rapid advancement of large language model (LLM) agents has raised new concerns regarding their safety and security, which cannot be addressed by traditional textual-harm-focused LLM guardrails. We propose GuardAgent, the first guardrail agent to protect other agents by checking whether the agent actions satisfy safety guard requests. Specifically, GuardAgent first analyzes the safety guard requests to generate a task plan, and then converts this plan into guardrail code for execution. In both steps, an LLM is utilized as the reasoning component, supplemented by in-context demonstrations retrieved from a memory module storing information from previous tasks. GuardAgent can understand different safety guard requests and provide reliable code-based guardrails with high flexibility and low operational overhead. In addition, we propose two novel benchmarks: EICU-AC benchmark to assess the access control for healthcare agents and Mind2Web-SC benchmark to evaluate the safety regulations for web agents. We show that GuardAgent effectively moderates the violation actions for two types of agents on these two benchmarks with over 98% and 83% guardrail accuracies, respectively.

Chat is not available.

Poster in Workshop: Workshop on Computer Use Agents

GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning

Zhen Xiang · Linzhi Zheng · Yanjie Li · Junyuan Hong · Qinbin Li · Han Xie · Jiawei Zhang · Zidi Xiong · Chulin Xie · Nathaniel Bastian · Carl Yang · Dawn Song · Bo Li

Poster
in
Workshop: Workshop on Computer Use Agents