יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

שיפור בזמן הריצה בעזרת דיפוזיה של דגימות שפה

Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
שיפור בזמן הריצה של דגימות שפה בעזרת דיפוזיה והדרכה על ידי שכר. פיתוח של דגימות שפה שמשתמשות בדיפוזיה ובהדרכה על ידי שכר, ומשפרות את זמן הריצה של דגימות שפה.
תקציר מקורי באנגליתarXiv:2602.22871v2 Announce Type: replace Abstract: Reasoning with large language models often benefits from generating multiple chains-of-thought, but existing aggregation strategies are typically trajectory-level (e.g., selecting the best trace or voting on the final answer), discarding useful intermediate work from partial or "nearly correct" attempts. We propose Stitching Noisy Diffusion Thoughts, a self-consistency framework that turns cheap diffusion-sampled reasoning into a reusable pool of step-level candidates. Given a problem, we (i) sample many diverse, low-cost reasoning trajectories using a masked diffusion language model, (ii) score every intermediate step with an off-the-shelf process reward model (PRM), and (iii) stitch these highest-quality steps across trajectories into a
קרא במקור המקורי