כתבה
arXiv cs.AI ·
מאמץ קשרי תפיסה לבסיס גרונדינג: בדיקת התערבות של נאמנות רשת
From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness
בדיקת התערבות של נאמנות רשת במודלי שפה גדולים. המחקר מבקר את נאמנות רשת המודלים ומציע דרך חדשה לבדיקת נאמנות רשת.
תקציר מקורי באנגליתarXiv:2609.23065v3 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input--output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as internal concept grounding: Does a large language model's (LLM) CoT reasoning engage the same internal concepts that support the LLM's direct prediction, and do the shared concepts causally drive its answer? Encoding a prediction pass and a CoT pass with a single shared sparse autoencoder (SAE), a reliable approximator of the latent concepts LLMs use, makes their internal concepts directly comparable. We introduce three correlational metrics of concept-level ali
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית