Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions

arXiv:2608.14642v1 Announce Type: new Abstract: Reinforcement Learning (RL) agents trained on a single reward signal exploit the gap between the designed reward and the intended behavior. This is particularly a problem when we are trying to imbue ethical behavior into RL agents. An agent can look ethical on average while concentrating its violations in a few bad episodes, and a creature in the environment harmed in one episode is not restored by good conduct in another. We compare four ways of t...

arXiv cs.LG ·Prabhjyot Singh, Majid Ghasemi, Mark Crowley ·
compartilhar: