יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אופטימיזציה ישירה של גיוון למסלולים מוצלחים

Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training
חוקרים פיתחו שיטה חדשה לאופטימיזציה של מודלים ללמידת מכונה, הנקראת Direct Diversity Optimization. השיטה משפרת את היכולת של המודלים למצוא מסלולים מוצלחים שונים. היא משלבת אלגוריתמים חדשים כדי לשפר את הגיוון והיעילות של המודלים.
תקציר מקורי באנגליתarXiv:2609.10052v1 Announce Type: cross Abstract: LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study this problem as successful strategy coverage: how broadly a model realizes distinct successful strategies under a fixed rollout budget. We present Direct Diversity Optimization (DDO), an offline post-training method that combines Divergence-Tree Collection (DTC) with the Reference-Relative Target-Odds Objective (RTO). DTC constructs state-aligned branch sets rooted at shared decision states, and RTO trains the model to match reference-relative targets over successful alternatives. DDO achieves the strongest task succ
קרא במקור המקורי