Blog
Robótica & RL
Risk-Aware Preference Learning for Stochastic Outcomes
arXiv:2607.15483v1 Announce Type: new Abstract: Learning reward functions from human preferences is a widely used approach for aligning robot behavior with user expectations in human-robot interaction. Most existing approaches assume that humans evaluate uncertain outcomes using expected utility (EU), aggregating outcome utilities linearly with their probabilities. However, behavioral evidence shows that humans are systematically risk-sensitive, overweighting rare negative events and exhibiting ...
arXiv cs.RO
·Yi-Shiuan Tung, Yuni Wu, Wei Jiang, Alessandro Roncone, Bradley Hayes
·
// relacionados
Leia também
Editorial
RynnValue: o relógio do vídeo como recompensa para robôs
Blog
Anthropic aplica marca d'água a todas as saídas do Claude globalmente, com marcas que "podem persistir mesmo após alguma edição"
Blog
webAI lança TwIL-LM: uma família de modelos de lógica formal de 1,7B e 3B para autoformalização em hardware local
Blog