יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Jev כ-Q או מדיניות בלמידת חיזוק

Can Jev be Your Q or Policy in Reinforcement Learning?
Jev הוא מודל דציסיוני שיכול לשמש כ-Q או מדיניות בלמידת חיזוק. מחקר זה בודק את יכולתו של Jev לשפר את יעילות הדגימה והביצועים של למידת חיזוק.
תקציר מקורי באנגליתarXiv:2610.11692v1 Announce Type: new Abstract: Foundation models supply reinforcement learning (RL) with priors that mitigate its longstanding weaknesses in sample efficiency and transfer, but their token-by-token generation makes queries sequential and costly. Jev, a recently released decision model, generates nothing and returns calibrated, typed answers in a single forward pass. Existing work studies foundation models in RL either as models to be trained or as generators to be prompted, and Jev belongs to neither category, having so far served only as a black box in single domains. How well such a model decides on its own in RL environments, and how it can improve RL as a component of training, therefore remain unaddressed. To this end, in this paper we first examine the requirements t
קרא במקור המקורי