יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

דיוורסיטי נתוני מסלול ספציפי להכשרה של LLM

Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents
במאמר זה, המחברים חקרו דיוורסיטי נתוני מסלול ספציפי בהכשרה של LLM דרך RL. הם הציגו פרוטוקול חדש שמאפשר לשמור ולשפר דיוורסיטי נתוני מסלול בצורה ספציפית. התוצאות המחקר הראו שהפרוטוקול החדש משפר את דיוורסיטי המסלולים בעודו משמר יעילות בביצועי המשימה.
תקציר מקורי באנגליתarXiv:2609.38805v1 Announce Type: new Abstract: LLM agents often admit multiple high-quality solutions to the same task, differing in reasoning structure, tool-use pattern, or interaction trajectory. Yet existing notions of diversity in LLM post-training are mostly implicit, arising from general stochasticity and regularization mechanisms rather than explicitly targeting task-relevant behavioral variation. While such implicit diversity can be useful, it does not directly specify which forms of behavioral variation should be encouraged for a given task. In this work, we study explicit trajectory diversity in RL-based post-training for LLMs. Our key idea is to define diversity through user-specified, task-specific trajectory descriptors, which map each sampled trajectory to an interpretable
קרא במקור המקורי