RLPF: Reinforcement Learning from Performance Feedback for Code Generation
arXiv:2607.27271v1 Announce Type: new Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness. This leaves an important gap for systems code: two programs can pass the same tests while differing greatly in runtime. We study how to train code agents to prefer faster correct implementations, rather than treating efficiency only as an evaluation metric. The key difficulty is that runtime is a fragile reward. It is meaningful only a...
arXiv cs.LG
·Huihao Jing, Haozhe Cui, Wenbin Hu, Shaojin Chen, Haochen Shi, Changxuan Fan, Yuxuan Liu, Hanyu Yang, Sirui Zhang, Ziyi Chen, Haoran Li, Yangqiu Song
·
// relacionados
Leia também
Blog
LingBot-Map Tutorial: GPU-Aware Inference and Point Cloud Export
Blog
Google Earth risked ruin with retracted AI tool for making fake satellite pics
Blog
Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids
Blog