יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידה לתכנן על ידי ראייה לאחור: היררכיות של היסטוריה לאחור להכשרת דגמי תקיפה

Learning to Plan by Looking Back: Hindsight Hierarchies for Training Reasoning Models
אנו מציגים שרשרת השכלה עצמית לדגמי תקיפה, המבוססת על ההבנה שפתרונות נוספים יכולים לאפשר לדגם לחקור רעיונות פתרון חשובים. השרשרת מתמקדת בהכשרת דגמי תקיפה להציג שלוש יכולות: חזות רעיונות פתרון מבעד לבעיות, הפיכת רעיונות פתרון לבעיות ופתרונות, ופתרון בעיות עם רעיונות פתרון. השרשרת מתמקדת בהכשרת דגמי תקיפה להציג שלוש יכולות: חזות רעיונות פתרון מבעד לבעיות, הפיכת רעיונות פתרון לבעיות ופתרונות, ופתרון בעיות עם רעיונות פתרון.
תקציר מקורי באנגליתarXiv:2610.12168v1 Announce Type: cross Abstract: We introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might enable the model to extract useful solution ideas in hindsight. We operationalize this by jointly training the same model to exhibit the following three capabilities: predicting solution ideas from problems alone, reverse-engineering ideas from problems and known solutions, and solving problems using provided ideas. The loop alternates between reverse engineering such ideas from problems with supplied solutions and using these ideas as additional supervision for joint training of all three capabilities. We give a formal speci
קרא במקור המקורי