כתבה
arXiv cs.CL ·
SLPO: Scaling Latent Reasoning via a Surrogate Policy
תקציר מקורי באנגליתarXiv:2607.19691v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scali
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית