כתבה
arXiv cs.CL ·
האמת הייתה תמיד נכונה: אליזה טובה בתרגילי נאמנות
The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes
במחקר חדש התגלה שאליזה טובה בתרגילי נאמנות עשויה להסתיר את האמת. זאת על ידי שימוש בפרובים לאמת, שמטרתם לגלות אם המודל נאמן או לא. אליזה טובה יכולה להסתיר את האמת על ידי שימוש באליזה שונה, שמפריך את הפרוב.
תקציר מקורי באנגליתarXiv:2609.10739v2 Announce Type: replace-cross Abstract: Linear probes that decode the truth from a language model's activations have been proposed as deception monitors, and a truth probe scoring below chance on a model trained to deceive is naturally read as evidence that the model hid or stopped representing the truth. We show that a fitting-label ambiguity can produce the same readout. In a controlled game, a model should report a secret bit to an ally and its complement to a rival. On compliant (ally) contexts the true bit and the answer the task prescribes are identical labels, so a probe fitted there cannot tell which of the two it measures. We call this complete agreement perfect aliasing. The labels are complements on rival contexts, so one ally-fitted probe scored against each h
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית