Inverse RL Helps Align AI by Imitating Humans
arXiv:2607.24900v1 Announce Type: new Abstract: Language model alignment aims to make model behavior reliably reflect desirable properties such as helpfulness, safety, and instruction following. Current approaches typically use supervised fine-tuning on demonstrations or reinforcement learning with rewards derived from verifiers or human feedback. These paradigms leave an important question underexplored: can demonstrations alone yield an implicit reward that can be inspected, reused, and optimi...
arXiv cs.LG
·Micha{\l} Wili\'nski, Liu Leqi, Chirag Nagpal
·