כתבה
arXiv cs.LG ·
BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards
תקציר מקורי באנגליתarXiv:2510.09596v2 Announce Type: replace Abstract: Today's generative models thrive with large amounts of supervised data and informative reward functions characterizing the quality of the generation. They work under the assumptions that the supervised data provides knowledge to pre-train the model, and the reward function provides dense information about how to further improve the generation quality and correctness. However, in the hardest instances of important problems, two problems arise: (1) the base generative model attains a near-zero reward signal, and (2) calls to the reward oracle are expensive. This setting poses a fundamentally different learning challenge than standard reward-based post-training. To address this, we propose BaNEL (Bayesian Negative Evidence Learning), an algo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית