Position: RL Researchers Need to Distinguish Between Solving Simulators and Using Simulators as a Proxy
arXiv:2606.28433v1 Announce Type: new Abstract: One goal in reinforcement learning (RL) research is to understand general-purpose sequential decision-making, using benchmark simulators as a proxy for learning in deployment settings. When running experiments, however, the goal of achieving high performance in the simulator can mutate into focusing exclusively on solving the simulator. To achieve high scores, researchers may adopt solutions exclusively meant for solving simulators, rather than lea...
arXiv cs.LG
·Matthew Vandergrift, Esraa Elelimy, Martha White
·
// relacionados
Leia também
Editorial
RynnValue: o relógio do vídeo como recompensa para robôs
Blog
Anthropic aplica marca d'água a todas as saídas do Claude globalmente, com marcas que "podem persistir mesmo após alguma edição"
Blog
webAI lança TwIL-LM: uma família de modelos de lógica formal de 1,7B e 3B para autoformalização em hardware local
Blog