יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

S^3-R1: למידה לחיפוש ולעניין בצעדים רב-שלביים עם נתונים סינתטיים

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data
S^3-R1 מציגה פלטפורמה שמשתמשת בנתונים סינתטיים ובאיתות רווחים צפופים כדי ללמוד חיפוש ועניין בצעדים רב-שלביים. הפלטפורמה נבחנה במבחן של 10% השיפור בכלליות רובוסטית על נתוני חוץ-תחום.
תקציר מקורי באנגליתarXiv:2605.01248v4 Announce Type: replace Abstract: Reinforcement learning (RL) post-training has enabled newer capabilities in models, such as agentic tool-use for search. However, these models struggle primarily due to limitations with sparse outcome-based rewards and a lack of training data that encapsulates questions of differing hardness, which results in models not performing deeper searches with tools to collect evidence for question-answering. To address these limitations, we introduce S^3-R1 (Synthetic data and stabilized Search R1), a framework that couples a data-centric approach with denser learning signals. We first develop a synthetic generation and curation pipeline that programmatically derives diverse, multi-hop questions from existing documents. This pipeline incorporates
קרא במקור המקורי