יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בדיקת גבולות כיווני אמת במודלים גדולים

Testing the Limits of Truth Directions in LLMs
חוקרים גילו גבולות לאוניברסליות של כיווני אמת במודלים גדולים. התוצאות מראות שכיווני אמת תלויים בשכבות המודל, סוג המשימה ורמת הקושי. המחקר מערער על ההנחה שכיווני אמת הם אוניברסליים.
תקציר מקורי באנגליתarXiv:2604.03754v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been shown to encode truth of statements in their activation space along a linear truth direction. Previous studies have argued that these directions are universal in certain aspects, while more recent work has questioned this conclusion drawing on limited generalization across some settings. In this work, we identify a number of limits of truth-direction universality that have not been previously understood. We first show that truth directions are highly layer-dependent, and that a full understanding of universality requires probing at many layers in the model. We then show that truth directions depend heavily on task type, emerging in earlier layers for factual and later layers for reasoning tasks
קרא במקור המקורי