Inverse RL Helps Align AI by Imitating Humans

arXiv:2607.24900v1 Announce Type: new Abstract: Language model alignment aims to make model behavior reliably reflect desirable properties such as helpfulness, safety, and instruction following. Current approaches typically use supervised fine-tuning on demonstrations or reinforcement learning with rewards derived from verifiers or human feedback. These paradigms leave an important question underexplored: can demonstrations alone yield an implicit reward that can be inspected, reused, and optimi...

arXiv cs.LG ·Micha{\l} Wili\'nski, Liu Leqi, Chirag Nagpal ·
compartilhar: