כתבה
arXiv cs.LG ·
Recirculation
תקציר מקורי באנגליתarXiv:2608.17981v3 Announce Type: replace Abstract: We describe an inference-time architectural enhancement for off-the-shelf foundation models that systematically reduces perplexity and boosts accuracy across generation and reasoning tasks. Our approach incurs essentially no additional latency during generation, though it requires serial processing in the prefill phase. Motivated by the fundamental limitation that state updates in feedforward transformers are bounded by model depth, our technique, recirculation, introduces a specific form of recurrence that allows the model to act as a dynamical system and track belief states. We distinguish this technique from chain-of-thought computation---which is better reserved for complex inferences rather than basic state tracking---as well as from
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית