כתבה
arXiv cs.AI ·
למידה של מה לאמן: ריפוי-מודד נתוני-עצמי לעבודת-עצמי של מודלי שפה
Learning What to Practice: Diagnosis-Guided Self-Evolution for Language Models
DiagEvo מגדיר את עבודת-עצמי של מודלי שפה על ידי היסטוריה של כשלים. המודל DiagEvo משתמש בהיסטוריה של כשליו כדי לפקח על יצירת שאלות. המודל DiagEvo נבדק ב-9 מבחנים והציג תוצאות טובות יותר מבסיסים. DiagEvo הציג תוצאות טובות יותר ב-9 מבחנים, כולל 72.3% דיוק ממוצע ב-5 מבחני סיבוכיות
תקציר מקורי באנגליתarXiv:2609.00768v2 Announce Type: replace Abstract: Self-play supports the self-evolution of language models, but solver performance can plateau or decline across rounds without guidance. Existing unguided methods typically use difficulty, learnability, or diversity signals to keep questions challenging and varied, without identifying which unresolved reasoning weaknesses to target. Existing guided methods rely on external task resources such as human examples, document corpora, or specified difficulty targets. We introduce DiagEvo, which guides question generation using the solver's failure history from self-play, without external task resources. Its diagnostician extracts recurring error causes and stores them in an error-cause memory. The memory groups related causes under skill nodes a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית