יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מתי ניתוחים פנימיים עוקפים קריאה לתשובה? קריאה לא-מדויקת וידע נסתר במודלי שפה

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models
ניתוחים פנימיים עוקפים קריאה לתשובה במודלי שפה בתנאים מסוימים. המחקר חושף קריאה לא-מדויקת וידע נסתר במודלי שפה. התגלית יכולה לשפר את יעילות המודלים ולאפשר פיתוח חדש של טכנולוגיות AI.
תקציר מקורי באנגליתarXiv:2609.04582v1 Announce Type: cross Abstract: A 0.6B language model, asked to verify 1,200 logical conclusions (half valid, half corrupted by a single semantic edit), answers YES every time. Judged by behavior it discriminates nothing; linear probes on its hidden states read the correct verdict at 0.96 AUC, transferring to unseen logical structures and separating foils built from exactly the words of the true conclusion (0.90). We ask where the verdict is lost, and find the dominant failure is a single scalar. The verdict survives to the model's own output logits (margin AUC 0.89) along a well-aligned readout direction; a saturated decision threshold, offset by +4.6 sigma, erases it. The diagnosis generalizes: across 90 semantic-label configurations of a five-model, three-family factor
קרא במקור המקורי