יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בדיקת ההנחה: בדיקות תקינות להשוואות RL עם נתונים ישנים במודלי שפה

Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models
במאמר זה, המחברים מציגים את PTH (Probe The Harness), סט של בדיקות שמאפשרות להפיק את ההנחה. הם מציגים דוגמאות של כיצד PTH יכול לשנות את התוצאות של השוואות RL. המחברים גם מציגים את ה-PPO ratio, ואת ההבדלים בין TIS ו-SAN.
תקציר מקורי באנגליתarXiv:2610.02911v1 Announce Type: cross Abstract: Methods for training language models on stale samples are judged by comparisons against importance-corrected baselines. We show that details of the experimental harness can reverse the observed ranking of methods, and we introduce PTH (Probe The Harness), a set of checks that makes the harness visible. Our case is a comparison between SAN, a behaviour-free method, and truncated importance sampling (TIS) on verl and in a single-GPU trainer, in which SAN first finished ahead in both stacks. Four details of the harness changed this comparison: the PPO ratio was taken against the learner's own recomputed probabilities, the data seed did not reach the TIS arm, the replay queue reused its first batch for 33 updates, and two loss normalisers diffe
קרא במקור המקורי