יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אימון רפורמנטי על תפוקות נבואיות לרגרסיה של LLM

Reinforcement Learning over Predictive Distributions for LLM Regression
אימון רפורמנטי לרגרסיה של LLM עם תפוקות נבואיות. המחקר מציג חדשנות באימון LLM לרגרסיה, עם תוצאות טובות יותר בקליברציה ובירידה בטעות הערכה.
תקציר מקורי באנגליתarXiv:2605.20740v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as flexible regressors capable of predicting real-valued quantities from heterogeneous inputs. Yet most LLM regression objectives optimize predictions independently, often yielding poor calibration. We introduce Distribution-Aware Reward (DAR), an on-policy reinforcement learning objective that instead jointly evaluates the empirical predictive distribution formed by multiple predictions for the same input. To translate this distribution-level objective into rollout-level rewards, we assign each prediction credit based on its leave-one-out contribution to the quality of the overall predictive distribution. This encourages predictions that are well-centered and appropriately dispersed around
קרא במקור המקורי