יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

לימוד לפתור בעיות קשות ב-RL ל-LLMs

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
שיטה חדשה ללימוד מחשבי מבוססת RL משפרת ביצועים של LLMs בבעיות קשות. השיטה, Never Give Up, משתמשת בדגימה אדפטיבית כדי להקצות יותר חישוב לבעיות קשות. השיטה הוכחה כיעילה בבנק אותות Deepscaler ובמשימת קוד Manufactoria.
תקציר מקורי באנגליתarXiv:2609.13443v1 Announce Type: cross Abstract: We demonstrate that training LLMs with RL does not improve performance equally across a dataset. RL shows large improvements on easy problems that an LLM is already good at solving, but small improvements on hard problems. We call this the Matthew Effect in RL for LLMs, after the phenomenon of cumulative advantage from economics and network science summarized as "the rich get richer". The naive explanation is that hard problems require more compute to find a solution. We argue that modern RL methods are exacerbating the issue by wasting too much compute on easy problems and instead should dynamically reallocate how they use compute. We introduce Never Give Up (NGU), a simple adaptive sampling method that keeps generating samples for a probl
קרא במקור המקורי