כתבה
arXiv cs.LG ·
בדיקת החניכה: בדיקות הגדרה להשוואות RL עם נתונים מיושנים במודלי שפה
Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models
חוקרים הציגו PTH, סדרת בדיקות להערכת השוואות RL עם נתונים מיושנים במודלי שפה. המחקר משווה בין שיטות SAN ו-TIS, ומציג פרטים על השפעת הפרטים הניסויים על התוצאות. הממצאים מראים כי TIS מתעלה על SAN במודל verl.
תקציר מקורי באנגליתarXiv:2610.02911v1 Announce Type: new Abstract: Methods for training language models on stale samples are judged by comparisons against importance-corrected baselines. We show that details of the experimental harness can reverse the observed ranking of methods, and we introduce PTH (Probe The Harness), a set of checks that makes the harness visible. Our case is a comparison between SAN, a behaviour-free method, and truncated importance sampling (TIS) on verl and in a single-GPU trainer, in which SAN first finished ahead in both stacks. Four details of the harness changed this comparison: the PPO ratio was taken against the learner's own recomputed probabilities, the data seed did not reach the TIS arm, the replay queue reused its first batch for 33 updates, and two loss normalisers differe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית