יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הפער השליטה: מה נמנעים משמירות LLM, ולמה לא יכולת

The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability
שמירות בטיחות LLM נכשלות בסיכול תכונות מסוימות, ולא יכולת. המאמר עוסק בפער השליטה של שמירות LLM, ובכך שהן לא יכולות לסכל תכונות מסוימות. המאמר כולל ניתוח של תכונות שמירות LLM, ובכך שהן לא יכולות לסכל תכונות מסוימות. המאמר גם כולל דיון בכך ששמירות LLM צריכות להיות עצמאיות, ולא תלויות במודלים או בכלים.
תקציר מקורי באנגליתarXiv:2609.07162v1 Announce Type: new Abstract: Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hyperproperties, witnessed only by two executions. The standard consequence is a binary impossibility: one trace cannot decide them. We replace the binary with a measurement. A tight bound puts the balanced accuracy of any single-trace monitor at $\tfrac12+\tfrac12\,TV(P_0,P_1)$, turning undecidability into a graded detectability frontier and defining an oversight gap: a monitor's shortfall below it. On a leak family with closed-form $TV$, nine LLM monitors are optimal at $TV=0$ but capture little signal as $TV$ grows; at $TV=1$, where a 20-line membership check scores $100\%$, they average $60.9
קרא במקור המקורי