כתבה
arXiv cs.CL ·
StateTree: שיפור תהליכי שיחה ארוכי טווח
StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning
StateTree הוא שיטה לשיפור תהליכי שיחה ארוכי טווח באמצעות למידת חיזוק. היא מאפשרת למודלים לנווט ברשומות מרובות ולפתור שאילתות מורכבות. StateTree עולה על מודלים אחרים כמו QwenLong-L1-32B
תקציר מקורי באנגליתarXiv:2609.38809v1 Announce Type: new Abstract: Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית