כתבה
arXiv cs.LG ·
VETTA: קואורדינציה של תרגיל ותווים להצמדת זכות ל-LLM Agents
VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents
VETTA מציגה שיטה ללמידת הצמדת זכות בשני רמות: תרגיל ותווים. השיטה, המכונה VETTA, משתמשת בקריטיקה קצרה ומשותפת ללמידת הצמדת זכות בשני רמות. VETTA נבחנה בשני מבחנים קשים, ALFWorld ו-WebShop, והציגה שיפורים בהצלחה של 37.5% ו-22.3%, בהתאמה.
תקציר מקורי באנגליתarXiv:2610.08402v1 Announce Type: new Abstract: Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two related credit-assignment questions: which responses helped achieve the outcome, and which generation decisions mattered within each response? Existing methods typically focus on only one level: turn-level methods evaluate complete responses but do not distinguish the decisions within them; token-level methods can propagate feedback across turns but do not explicitly model credit for each response. These complementary limitations motivate learning credit at both levels and coordinating it in a single policy update. We introduce VETTA, a credit assignment method that jointly learns turn- and toke
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית