כתבה
arXiv cs.LG ·
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
תקציר מקורי באנגליתarXiv:2607.14952v3 Announce Type: replace Abstract: Long-context RL post-training is constrained by the lifetime of state and gradients, not attention cost alone. In GRPO, one multi-million-token prompt must serve old-policy and reference scoring plus multiple policy responses, while conventional autograd keeps the prompt graph and all response graphs live alongside model weights, caches, and distributed communication buffers. We present LongStraw, an objective-aware, architecture-aware system for resident-state virtualization, response replay, and distributed-gradient execution. Its transaction captures the shared prompt without autograd, retains only the architecture-required state on explicitly owned pages, restores that state for each group member, scores old/reference branches without
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית