יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מעקב אחר גבולות ניידים

Tracking the Moving Frontier: Long-Short Term Advantage Estimator
LSTAE הוא אלגוריתם RL יעיל המשתמש בניסיון היסטורי להערכת יתרונות. הוא משמר מעקב קבוע לכל מטלה ומודד את התרומה היחסית של כל טראקטוריה חדשה.
תקציר מקורי באנגליתarXiv:2609.06671v1 Announce Type: new Abstract: Group-based RLVR methods estimate advantages by repeatedly sampling multiple trajectories for each prompt, making long-horizon agent training expensive and discarding useful experience accumulated across iterations. We ask whether historical experience can replace these repeated within-iteration comparisons without directly optimizing on stale trajectories. We introduce Long-Short Term Advantage Estimator (LSTAE), a single-stream RL algorithm that uses history for advantage estimation while updating the policy only with the current rollout. LSTAE maintains a persistent tracker for each task anchor. At the trajectory level (long term), a drift-aware historical baseline tracks the anchor's moving success frontier and measures the relative contr
קרא במקור המקורי