יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

OpenJev-RLCD: יישום RLCD עובד

OpenJev-RLCD: A Working RLCD Implementation
OpenJev-RLCD הוא יישום עובד של RLCD, שיטה לקבלת החלטות מתוך דגמים. המודל מנבא רציונל ומדרג את התפלגות התשובות. Qwen3-1.7B הראה תוצאות טובות במשימות ספציפיות.
תקציר מקורי באנגליתarXiv:2609.38850v1 Announce Type: cross Abstract: Decision models such as Jev answer questions with probabilities, which are only useful if they are calibrated. Open-source reproductions rely on supervised fine-tuning plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models: the model samples a rationale, and we score the answer distribution it commits to afterwards with a strictly proper scoring rule. A variance identity shows that scoring the mixture of several samples rewards disagreeing rationales, and that RLVR is exactly this mixture objective without its diversity term. Optimized naively, the per-rationa
קרא במקור המקורי