Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
arXiv:2607.07405v1 Announce Type: new Abstract: Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive environments, a tool may execute any well-formed call even when the corresponding state transition is forbidden by domain policy. The result is a silent wrong state (a booking cancelled, a passenger count changed, a claim acted on without verification) that neither the tool nor the agent's self-rep...
arXiv cs.AI
·Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu
·
// relacionados
Leia também
Blog
As alegações mais escandalosas no processo da Apple contra a OpenAI por segredos comerciais
Blog
O que a mais recente descoberta em IA da Anthropic mostra — e o que não mostra
Blog
Novo guia de prompts da OpenAI diz aos usuários para parar de complicar e começar pelo resultado
Blog