יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת סיגנלים בין-אגנטים מאפשרת הכשרה מתואמת של LLM לפרקים

Cross-Agent Learning Signals Enable Coordinated Role-Decomposed LLM Training
מאמר חדש מציג פרקטיקה להכשרת LLM לפרקים, על ידי שימוש בסיגנלים למידה בין-אגנטים. הפרקטיקה, DAC, מאפשרת ל-LLM לבחור אם להפיק תשובה או להפסיק, ולקבוע זכויות לכל פרק. המאמר מציג תוצאות של DAC על-פני שבעה בנקאי QA ושני רכיבי קוד. DAC נראה כמוצלח יותר מבסיסים חזקים של LLM ו-RAG.
תקציר מקורי באנגליתarXiv:2606.10684v2 Announce Type: replace Abstract: Agentic search systems must coordinate evidence acquisition and response generation, yet existing approaches either couple both roles under a single agent objective or decompose them without disentangling their respective contributions to the final outcome. We introduce DAC (Divide and Cooperate), a role-decomposed training framework that, given task-specific external verification signals, trains a searcher and a generator with role-specific cross-verification rewards. DAC allows the generator to abstain when the retrieved evidence appears insufficient, and uses this decision together with externally evaluated search sufficiency to assign appropriate credit to each role. To prevent degenerate over-abstention, we further introduce hard-pos
קרא במקור המקורי