יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

DART-ES: שיטה חדשה לקיפול LLMs

DART-ES: Difficulty-Aware Reweighting and Targeted Replay for Fine-Tuning LLMs with Evolution Strategies
DART-ES היא שיטה חדשה לקיפול מודלי שפה גדולים (LLMs) באמצעות אסטרטגיות אבולוציה. השיטה משפרת את ביצועי הקיפול ומפחיתה את הזיכרון וזמן הריצה. DART-ES עובדת עם מודלים שונים, כולל LLaMA.
תקציר מקורי באנגליתarXiv:2610.06993v1 Announce Type: cross Abstract: Evolution Strategies (ES) enable memory efficient full parameter fine-tuning of large language models (LLMs) using only forward computation. However, standard ES uniformly averages rewards across problems and compresses problem level population feedback into a single scalar, making it difficult to capture how the learning value of each problem changes with model capability. To address this limitation, we propose Difficulty-Aware Reweighting and Targeted Replay for Evolution Strategies (DART-ES). DART-ES estimates the local solvability of each problem from its pass rate across the perturbation population and aggregates historical observations to construct a dynamic difficulty state. This shared state jointly guides continuous difficulty rewe
קרא במקור המקורי