Blog
Robótica & RL
P3: Probabilistic Policy Propagation for Stable VAE-Based Robot Learning
arXiv:2607.25541v1 Announce Type: new Abstract: Variational Autoencoders are widely used to encode high-dimensional and noisy observations in robotics. However, their stochastic latent creates a mismatch with Proximal Policy Optimization (PPO): an effective policy marginalizes over the latent distribution, whereas former implementations estimate its probability ratio and KL divergence using only one latent sample. We identify a fundamental but overlooked theoretical cause: naive single-sample ap...
arXiv cs.RO
·Liyun Yan, Jianming Ma, Yang Zhang, Shengcheng Fu, Zhanxiang Cao, Keqi Zhu, Yizhi Chen, Yue Gao
·
// relacionados