PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

arXiv:2608.17289v1 Announce Type: new Abstract: Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectories even when these trajectories differ substantially in their interaction efficiency. For instance, circuitous successes are often assigned the identical outcome reward, causing advantage collapse and severe perfor...

arXiv cs.AI ·Dayang Liang, Liyuan He, Xuan Feng, Shuxin Li, Bo An, Yunlong Liu ·
compartilhar: