Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity

arXiv:2607.28849v1 Announce Type: new Abstract: Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories of problems, such as meta-learning, hierarchical task decomposition, and reinforcement learning from human feedback (RL-HF). Most of the bilevel RL algorithms are either not scalable because of using hypergradient with Hessian, or they suffer from high sample complexity because of using penalty-based approximati...

arXiv cs.LG ·Naman Saxena, Mudit Gaur, Vaneet Aggarwal ·
compartilhar: