יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

DRACO: שיטה לאימון סוכנים באופן עדין

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
DRACO היא שיטה חדשה לאימון סוכנים באמצעות רובריקות דינאמיות. היא מאפשרת אימון סוכנים ללא צורך באישורים. DRACO זכתה לתוצאות טובות בניסויים.
תקציר מקורי באנגליתarXiv:2609.04094v1 Announce Type: cross Abstract: Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored once per trajectory, but a single scalar is a poor signal across tens of steps. We propose DRACO: Distributing Rubric-based Advantage for Credit Optimization. It generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages i
קרא במקור המקורי