Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgmen...
arXiv cs.AI
·Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, Bowen Liu
·