כתבה
arXiv cs.LG ·
DART: אימון מחזורי עם התפלגות אדוורסרית
DART: Distributional Adversarial Recurrent Training for Algorithm Learning
DART הוא שיטת אימון חדשה המשתמשת בהתפלגות אדוורסרית מקומית סביב הפתרון הנכון. היא משפרת את איכות הפתרון, יציבות ועמידות של מודלים מחזוריים.
תקציר מקורי באנגליתarXiv:2609.05988v1 Announce Type: new Abstract: Recurrent reasoning models (RRMs) can solve structured problems, achieving easy-to-hard generalization through iterative computation in hidden space. These models are typically trained with instance-level supervision, which becomes increasingly problematic as task difficulty grows: valid solutions occupy a tiny region of the solution space, while invalid solutions proliferate rapidly. We propose Distributional Adversarial Recurrent Training (DART), a training framework that replaces single-point supervision with a local target distribution around the ground-truth solution and aligns model outputs with this distribution through an adversarial objective. DART provides a richer learning signal and encourages more stable iterative trajectories to
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית