כתבה
arXiv cs.LG ·
שיפור ניטרול סקאלי עם משמרי נלמדים
Improving scalable oversight with co-trained monitors
ניתן לשפר את השליטה הסקאלית בעזרת משמרים שנלמדים ביחד עם העובדים.
תקציר מקורי באנגליתarXiv:2609.36049v1 Announce Type: new Abstract: Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study whether this failure mode can be avoided by co-training the monitor alongside the worker, and explore both supervised and self-supervised approaches. In the supervised setting, we prove a characterization: monitoring is possible with vanishing error and query rates exactly when the class of possible monitor functions has finite Littlestone dimension. This connects worker monitoring with an established literature on adversarial online learning. For self-supervision, we propose a co-training procedure based on test-time distillation: the monitor uses additional test-time compute to generate traini
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית