יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

VERPO: אופטימיזציה של מדיניות עם הוכחות מאומתות

VERPO: Verified Evidence Regularized Policy Optimization
VERPO היא שיטה חדשה לאופטימיזציה של מדיניות, המשתמשת בהוכחות מאומתות כדי לשפר את ביצועי המודל. השיטה נבדקה על מודלים Qwen ו-Llama, והראתה שיפורים משמעותיים בתוצאות.
תקציר מקורי באנגליתarXiv:2609.06100v1 Announce Type: new Abstract: Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-conditioned Teachers provide denser supervision by replaying sampled trajectories with privileged feedback. Yet indiscriminate imitation risks transferring formatting or reasoning-style shifts that do not support task success. We introduce VERPO, a Verified Evidence Regularized Policy Optimization framework that treats evidence as a proposal for policy correction while retaining the outcome objective. It separates evidence-free reference restoration from signed token-level evidence corrections. Fisher Evidence Contrast attenuates corrections along an estimated evidenc
קרא במקור המקורי