כתבה
arXiv cs.LG ·
EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
תקציר מקורי באנגליתarXiv:2606.18967v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a representative post-training paradigm for large language models (LLMs), enabling strong reasoning and agentic capabilities. However, rollout generation remains a dominant latency bottleneck because autoregressive (AR) sampling decodes responses sequentially and a small number of long-tailed generations often determine completion time. Speculative decoding (SD) is a well-established technique for serving fixed LLMs that reduces latency by rapidly drafting tokens and accepting them through parallel verification while preserving the target-model distribution. However, its practical speedups do not directly carry over to RL rollouts: (i) the evolving target policy makes any fixed drafter increasingly m
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית