כתבה
arXiv cs.CL ·
גילוי רמאות מחוץ לתחום האימון
Probe Generalization as Subspace Selection for OOD Deception Detection
חוקרים מצאו שניתן לשפר את יכולת הגילוי של רמאות מחוץ לתחום האימון על ידי שימוש ב-Llama-3.1-8B-Instruct. הם השתמשו בשיטת סלקציה של תת-מרחב כדי לשפר את היכולת לזהות רמאות.
תקציר מקורי באנגליתarXiv:2609.02893v1 Announce Type: new Abstract: Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes over 3 held-out deception detection datasets, we find that projecting inputs onto a small subset of principal components (PCs) from the training distribution of activations enables cross-domain transfer that nearly matches the performance of probes trained directly on the test distribution. Furthermore, we find that PC interpretations can be used to find a subset of those transferable PCs. By using an LLM judge to score each PC on whether its most/ least activating examples imply a transferable deception direction, th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית