כתבה
arXiv cs.AI ·
Co-Evolving LLM Evaluators and Policies via DynamicRubric
במאמר זה, המחברים מציגים פרקטיקה חדשה לשיפור LLMs דרך פידבק של מבחנים ואופטימיזציה של מדדי פוליצי. הם מציגים את DynamicRubric, פרקטיקה שמפקידה רוביקים דינמיים לכל קבוצת משיבים ומסכמת את הדינמיקה שלהם. המחברים מדגימים את הפרקטיקה שלהם באמצעות ניסויים עם LLMs של 8B ו-70B, ומציגים תוצאות טובות בשיפור מדדי פוליצי ובשיפור בביצועים של LLMs.
תקציר מקורי באנגליתarXiv:2607.20083v2 Announce Type: replace-cross Abstract: Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become close in quality. These close candidates create a bottleneck for policy optimization: collapsed relative evaluator score gaps yield weak or misleading policy supervision. We theoretically characterize why these gaps matter through a probability allocation view, showing that the directional gain of shifting probability mass from one response to another is exactly the evaluator score gap between them. This identifies relative score gaps as the policy optimization signals that guide updates. Motivated by this view, we propose DynamicRubric, a response-set-co
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית