A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

arXiv:2607.25207v1 Announce Type: new Abstract: This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online interactions with a target environment and offline data from a source environment. A central challenge is that offline data may be collected from outdated environments with shifted transition dynamics, making naive integration of historical data ineffective. To address this, ...

arXiv cs.LG ·Zheshun Wu, Renjie Zheng, Jinhang Zuo, Zenglin Xu, Fang Kong ·
compartilhar: