יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

UniRRM: דגם פרסונלי של תגמולים תקין למודלי רכיבה

UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms
UniRRM היא דגם פרסונלי של תגמולים תקין שתומך במספר שפות ובמגוון פרדיגמות הערכה. היא משתמשת בשרשרת סיבובים ליצירת קריטריות תקינות ומדיניותיות.
תקציר מקורי באנגליתarXiv:2609.05910v1 Announce Type: cross Abstract: Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack interpretability. Recent works on generative reward models offer a promising alternative, but they remain constrained by static evaluation criteria, fragmented evaluation paradigms, and limited multilingual support. To address these challenges, we introduce \textbf{MixReward}, a large-scale multilingual dataset spanning six domains and 103 languages, containing both pairwise and listwise data, and propose \textbf{UniRRM}, a unified reasoning reward model supporting multiple
קרא במקור המקורי