יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אימון קריטי אגנטי

Agentic Critical Training
אימון קריטי אגנטי: שיטה חדשה לאימון מודלי שפה לשפוט פעולות. השיטה, המכונה ACT, משתמשת בלמידה עצמית עם שכר הניתן לאימות, ומאפשרת למודלים לפתח סיבות לפעולותיהם. ACT נבדקה במספר סביבות, כולל ALFWorld-ID, WebShop, ו-ScienceWorld, והראתה תוצאות טובות יותר משיטות אחרות. המחברים טוענים כי ACT יכולה לשפר את הביצועים של מודלי שפה בתחומים שונים, כולל הפקת טקסט, הפקת תמונות, ועוד.
תקציר מקורי באנגליתarXiv:2603.08706v2 Announce Type: replace Abstract: Imitation learning (IL) teaches language-model agents to reproduce expert actions but not to distinguish them from plausible mistakes. Self-reflection methods expose models to alternatives yet use supervised fine-tuning (SFT) to imitate fixed rationales and actions. We introduce Agentic Critical Training (ACT), which uses reinforcement learning with verifiable rewards (RLVR) to train models to judge actions directly. At each expert-trajectory state, ACT pairs an expert action with an alternative sampled from the initial policy and randomizes their order. The model generates its own reasoning but is rewarded only for selecting the expert action. ACT reuses demonstrations, requires no reference rationales, and allows pair reuse across model
קרא במקור המקורי