יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

OPSRD: תפיסת תפקיד עצמית ללא פתרונות

OPSRD: On-Policy Self-Role Distillation
תפיסת תפקיד מגדילה את התנהגות ספציפית של מודלי שפה גדולים. OPSRD מציעה דרך להפיץ את התפיסה עצמית, כולל תפקידים קבועים, כדי לשפר את הדיוק של המודל. ניתן להשתמש ב-OPSRD כדי לשפר את הביצועים של המודל במשימות שונות.
תקציר מקורי באנגליתarXiv:2609.39884v1 Announce Type: cross Abstract: Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond
קרא במקור המקורי