כתבה
arXiv cs.AI ·
Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents
תקציר מקורי באנגליתarXiv:2609.38805v1 Announce Type: cross Abstract: LLM agents often admit multiple high-quality solutions to the same task, differing in reasoning structure, tool-use pattern, or interaction trajectory. Yet existing notions of diversity in LLM post-training are mostly implicit, arising from general stochasticity and regularization mechanisms rather than explicitly targeting task-relevant behavioral variation. While such implicit diversity can be useful, it does not directly specify which forms of behavioral variation should be encouraged for a given task. In this work, we study explicit trajectory diversity in RL-based post-training for LLMs. Our key idea is to define diversity through user-specified, task-specific trajectory descriptors, which map each sampled trajectory to an interpretabl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית