כתבה
arXiv cs.LG ·
CompactionRL: רכיבת למידת רב-מודלים עם קיצור תצוגה לאגנטים באורך-זמן
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
CompactionRL היא שיטת למידת רב-מודלים שמשלבת קיצור תצוגה לאגנטים באורך-זמן. היא מאפשרת לאגנטים ללמוד ממסלולים ארוכים ולשפר את ביצועיהם.
תקציר מקורי באנגליתarXiv:2607.05378v2 Announce Type: replace Abstract: Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by summarizing previous interaction states and continuing the rollout under a compressed context, but incorporating compaction into reinforcement learning remains underexplored. We propose CompactionRL, a reinforcement learning strategy to train long-horizon agentic LLMs with context compaction. Our approach jointly optimizes task execution and summary generation with token-level loss normalization and cross-segment generalized advantage estimation. This design enables the LLM agents to learn from compacted long-hori
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית