יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

לקסיקוגרפיה: פוסט-אימון רב-מטרותי להטמעת פוליסים

Lexicographic Multi-Objective On-Policy Distillation
במאמר זה, המחברים מציגים שיטה חדשה ללמידת רב-מטרותי של דגמי תקשורת, המאפשרת לשמור על יכולות גבוהות של דגמי GPT-5 ו-Gemini.
תקציר מקורי באנגליתarXiv:2610.02359v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then lo
קרא במקור המקורי