יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אמינות חד-ממוצע: ניצני זיווג CoT-Monitor חושפים רגישות תלויה ברציונל

A False Average: Pooled CoT-Monitor Accuracy Conceals a Reasoning-Dependent Fragility
מפקחי CoT המספרים על ניצני זיווג נתגלו ככושלים, עקב אמינות חד-ממוצע. המאמר חושף רגישות תלויה ברציונל של המפקחים.
תקציר מקורי באנגליתarXiv:2608.00583v3 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) monitors are reported by their aggregate accuracy on a pool of reward hacks. We show that this number is a false average. On Terminal Wrench, about 77% of hacks are given away by the actions alone, and the monitor's pooled accuracy is dominated by them; on the remaining 23%, where the reasoning is the only signal, the same monitor is fragile. We expose the fragility with a controlled attack: we rewrite only the agent's reasoning to read as good-faith engineering, leaving every command and output byte-identical, so the exploit is unchanged. One gradient-free rewrite drops a held-out monitor's catch rate on that subset from about 95% to between 4 and 11%, while the pooled rate falls only about 25 points, the sub
קרא במקור המקורי