יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR

ThinkPrior מציע פריורים חסרי רולאוטים לבחירת הפרסומים בתחילת הפעילות ב-RVLR. הפריור נבנה בעזרת קצאנך חיצוני, ומאפשר לבחור את הפרסומים המועדפים ללא צורך ברולאוטים.
תקציר מקורי באנגליתarXiv:2609.09075v1 Announce Type: new Abstract: In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimization (GRPO), the KL-free reward-advantage term studied here depends on within-group reward variation. If all rollouts in a group are correct or all are wrong, their group-relative advantages are identically zero; these zero-advantage silent groups provide no reward-advantage gradient, yet uniform sampling spends 39% of a run's rollouts on them. History-based prompt selection must first spend target-policy rollouts to estimate difficulty, creating a cold start with rollout waste; ThinkPrior instead uses an external anchor in one offline pass to construct a zero-rollout difficulty prior before the first target-policy rollout. The verifier-scored
קרא במקור המקורי