Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups
arXiv:2608.17423v1 Announce Type: new Abstract: GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. This simplification comes with a sampling cost: group-relative advantages require multiple rollouts from each scene. Under binary success rewards, groups whose rollouts all succeed or all fail have zero advantage and are discarded by dynamic sampling. These groups are especially common early in tr...
arXiv cs.RO
·Zeyun Deng, Yuzhe Lu, Yawei Wang, Linbo Liu, Qing Ping, Han Ding, Guande Wu, Panpan Xu, Jun Huan
·
// relacionados
Leia também
Blog
Flight attendants freaked out that Google is buying tons of Spirit employee data
Blog
Attackers are using AI to build exploits for industrial control systems, U.S. agencies warn
Blog
AI labs are failing to keep their own systems in check
Editorial