כתבה
arXiv cs.CL ·
בדיקת עובדות לטנטית
Latent Fact-Checking: Detecting Misinformation through Activation Engineering
חוקרים פיתחו שיטה לזיהוי מידע כוזב באמצעות הנדסת אקטיבציה. השיטה משתמשת במודלים מתקדמים כמו Llama ו-Qwen.
תקציר מקורי באנגליתarXiv:2608.06417v4 Announce Type: replace-cross Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level linguistic features or external knowledge retrieval, we examine truthfulness as a geometric property of a language model's representation space. We introduce a misinformation detection framework grounded in activation engineering, which leverages the latent geometry of transformer models. Our approach elicits a misinformation direction in the residual stream by contrasting activations from paired truthful and false statements, following the difference-in-means principle of Contrastive Activation Addition (CAA). At inference time, the last-token activation of an unseen claim is projected on
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית