כתבה
arXiv cs.AI ·
האמת לא נעלמה: אליזה מושלמת בתפקוד-סביבה
The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes
נמצא כי תפקוד-סביבה מושלם יכול לגרום לכשל בדפוסי זיהוי שקרים.
תקציר מקורי באנגליתarXiv:2609.10739v2 Announce Type: replace-cross Abstract: Linear probes that decode the truth from a language model's activations have been proposed as deception monitors, and a truth probe scoring below chance on a model trained to deceive is naturally read as evidence that the model hid or stopped representing the truth. We show that a fitting-label ambiguity can produce the same readout. In a controlled game, a model should report a secret bit to an ally and its complement to a rival. On compliant (ally) contexts the true bit and the answer the task prescribes are identical labels, so a probe fitted there cannot tell which of the two it measures. We call this complete agreement perfect aliasing. The labels are complements on rival contexts, so one ally-fitted probe scored against each h
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית