כתבה
arXiv cs.AI ·
SPIRAL: למידת חיפוש ואיגוד
SPIRAL: Learning to Search and Aggregate
SPIRAL הוא כלי לשיפור תהליכי היגיון של מודלי שפה. הוא מאפשר למודלים להשתמש במספר רב של רצפים ולאגד אותם לתוצאה סופית. הניסויים הראו ש-SPIRAL משפר את ביצועי המודלים.
תקציר מקורי באנגליתarXiv:2606.23595v2 Announce Type: replace Abstract: Language model reasoning can be substantially improved at test time via scaffolds that scale inference compute across different primitives -- sequential reasoning within a trace, independently sampled parallel traces, and aggregation of multiple reasoning traces into a final response. During post-training, however, language models are optimized only for sequential reasoning within a single trace. We introduce Sequential-Parallel-Aggregative Reinforcement Learning (SPIRAL), a framework in which a language model is trained to use all three primitives, as part of a unified inference compute pipeline. Concretely, the language model first samples a set of independent traces in parallel, each produced through sequential chain-of-thought reasoni
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית