Blog
LLMs & Texto
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
arXiv:2607.26120v1 Announce Type: new Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In these settings, misalignment with collective goals becomes a central concern. We propose a novel framework for evaluating objective misalignment using the social deduction game Werewolf, modifying the objective of a singl...
arXiv cs.AI
·Marylou Fauchard, Florian Carichon, Margarida Carvalho, Golnoosh Farnadi
·
// relacionados