יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חקירה יעילה לאופטימיזציה של נאש

Efficient Exploration for Iterative Nash Preference Optimization
חוקרים פיתחו שיטה חדשה לאופטימיזציה של מודלי שפה גדולים, באמצעות חקירה יעילה ל-Nash. השיטה, הנקראת ENPO, משלבת רגולריזציה עם חקירה אדוורסרית. הניסויים הראו שיפורים משמעותיים בהשוואה לשיטות קיימות, עם מודל Llama-3-8B-Instruct.
תקציר מקורי באנגליתarXiv:2606.01382v2 Announce Type: replace-cross Abstract: Preference alignment is central to improving large language models (LLMs), but reward-based formulations can be restrictive when human preferences are non-transitive. Nash learning from human feedback (NLHF) addresses this limitation by modeling alignment as a preference game and seeking a Nash equilibrium. However, the learning-theoretic foundations of scalable NLHF remain limited: existing regret guarantees rely on explicit preference-model estimation and minimax oracles, whereas simpler iterative methods lack such guarantees. We study online iterative NLHF and identify exploration as a key obstacle. First, we show that standard iterative NLHF can incur an exponential dependence on the inverse KL-regularization parameter, demonstr
קרא במקור המקורי