יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אימון שופטי LLM מבוססי חומרה על פידבק שפתי

Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
אימון שופטי LLM מבוססי חומרה על פידבק שפתי. המחקר חוקר את אימון שופטי LLM מבוססי חומרה על פידבק שפתי, ומציג טכניקה חדשה של עדיבה-בחירה-עדיבה (position-selective self-distillation) שמשפרת את הביצועים של שופטי LLM.
תקציר מקורי באנגליתarXiv:2609.38792v1 Announce Type: new Abstract: We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions
קרא במקור המקורי