יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

שיפור ניטרול סקאלי עם מפקחים מאומנים

Improving scalable oversight with co-trained monitors
אנו מחקרים את האפשרות לאבטחה סקאלית עם מפקחים שנלמדו ביחד עם העובדים. ניתן להגיע לאבטחה סקאלית עם מפקחים שנלמדו ביחד עם העובדים.
תקציר מקורי באנגליתarXiv:2609.36049v1 Announce Type: cross Abstract: Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study whether this failure mode can be avoided by co-training the monitor alongside the worker, and explore both supervised and self-supervised approaches. In the supervised setting, we prove a characterization: monitoring is possible with vanishing error and query rates exactly when the class of possible monitor functions has finite Littlestone dimension. This connects worker monitoring with an established literature on adversarial online learning. For self-supervision, we propose a co-training procedure based on test-time distillation: the monitor uses additional test-time compute to generate trai
קרא במקור המקורי