כתבה
arXiv cs.LG ·
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?
תקציר מקורי באנגליתarXiv:2607.27203v1 Announce Type: new Abstract: Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without needing to pretrain the Q-function. In this paper, we systematically study whether pretraining the Q-function actually helps when fine-tuning on top of a pretrained base policy. We find, surprisingly, that naive Q-function pretraining often provides little benefit over random initialization. We show this s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית