יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

עקיפת אמון: למה מנטורים לא מעבירים ב间 משפחות

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
חוקרים גילו כי מנטורים מהימנים לא מעבירים היטב בין משפחות שונות של מודלים. המחקר מראה כי הדיוק של המנטורים יורד בצורה משמעותית כאשר הם מועברים למשפחות אחרות, וכי יש צורך בדרכים חדשות להעריך את היעילות של המנטורים.
תקציר מקורי באנגליתarXiv:2607.06596v2 Announce Type: replace-cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicious are audited or deferred. Such monitors are evaluated against one or two untrusted models, and the accuracy is reported as a property of the monitor. We ask whether it is partly a property of the pairing. We make the untrusted policy family the controlled axis: we fit a monitor on family A's transcripts, apply it to family B, and decompose the cross-family AUROC into how obvious each family's sabotage is, how capable each monitor is, and the residual own-family advantage after both are removed: the interaction. On code-backdoor transcripts the interaction is positive and survives the
קרא במקור המקורי