כתבה
arXiv cs.AI ·
פיתוח תזמון בשלב ההסתכלות עם דגמי שפה דיפוזיים על ידי חיבור מודרך
Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
במאמר זה, המחברים מציגים פרקטיקה חדשה לשיפור ביצועי דגמי שפה דיפוזיים. הם מציעים פרוטוקול של חיבור מודרך, המאפשר לשפר את תזמון הביצועים של דגמי שפה דיפוזיים. הפרקטיקה כוללת שלושה שלבים: סימולציה של דגמי שפה, חיבור מודרך וביצוע של פעולות סופיות. המחברים מציגים תוצאות מבחנות שמציגות את יעילות הפרקטיקה.
תקציר מקורי באנגליתarXiv:2602.22871v2 Announce Type: replace-cross Abstract: Reasoning with large language models often benefits from generating multiple chains-of-thought, but existing aggregation strategies are typically trajectory-level (e.g., selecting the best trace or voting on the final answer), discarding useful intermediate work from partial or "nearly correct" attempts. We propose Stitching Noisy Diffusion Thoughts, a self-consistency framework that turns cheap diffusion-sampled reasoning into a reusable pool of step-level candidates. Given a problem, we (i) sample many diverse, low-cost reasoning trajectories using a masked diffusion language model, (ii) score every intermediate step with an off-the-shelf process reward model (PRM), and (iii) stitch these highest-quality steps across trajectories
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית