יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מתכון לתפיסה בהקשר ארוך במודלים של שפה גדולה

A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation
חוקרים פיתחו שיטה חדשה לאימון מודלים של שפה גדולה לתפיסה בהקשר ארוך. השיטה משלבת טכניקות שונות, כולל אופטימיזציה ודיסטילציה, כדי לשפר את היכולת של המודלים להבין טקסטים ארוכים. החוקרים בדקו את השיטה החדשה על מודל ה-LLaMA וגילו שהיא משפרת את התוצאות.
תקציר מקורי באנגליתarXiv:2605.12227v3 Announce Type: replace Abstract: Existing approaches to post-train models for long-context tasks face complementary limitations: (i) supervised fine-tuning (SFT) provides stable supervision but suffers from exposure bias; (ii) reinforcement learning methods such as Group Relative Policy Optimization (GRPO) train on model-generated trajectories but struggle with long-horizon credit assignment and sparse rewards; and (iii) on-policy distillation (OPD) provides dense token-level guidance but does not directly optimize task rewards. We study these complementary strategies for long-context alignment and derive a recipe that combines GRPO with OPD-style teacher guidance: the student learns from its own rollouts using outcome-level rewards, while a stronger teacher provides den
קרא במקור המקורי