כתבה
arXiv cs.CL ·
SERL-SQL: חידוש בהכווין חזרה לאחור לטקסט-SQL: הכווין חזרה לאחור נבחרת ללמידת רפורמציה
SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning
SERL-SQL: חידוש בהכווין חזרה לאחור לטקסט-SQL: הכווין חזרה לאחור נבחרת ללמידת רפורמציה. ניתן להכווין חזרה לאחור נבחרת ללמידת רפורמציה על ידי שימוש בטכניקת הכווין חזרה לאחור. המאמר עוסק בפיתוח טכניקה חדשה ללמידת רפורמציה, המבוססת על טכניקת הכווין חזרה לאחור. הטכניקה, הנקראת SERL-SQL, מאפשרת למערכת ללמוד לבחור בצורה נבחרת את הפעולות לביצוע.
תקציר מקורי באנגליתarXiv:2608.00485v4 Announce Type: replace Abstract: Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods use execution correctness only as a trajectory-level reward, which provides limited guidance for identifying the SQL decisions responsible for success or failure. We propose SERL-SQL, a selective execution-grounded reinforcement learning framework for multi-turn Text-to-SQL agents. SERL-SQL samples on-policy SQL interaction trajectories and uses a training-only teacher to re-score student actions with execution feedback. The resulting teacher--student likelihood gap is converted into bounded, masked weights that reweight GRPO advantages only on SQL and tool-action tokens. In this way, task
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית