כתבה
arXiv cs.LG ·
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
תקציר מקורי באנגליתarXiv:2602.02639v2 Announce Type: replace-cross Abstract: LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithfulness metrics have critical limitations, typically relying on identifying unfaithfulness via adversarial prompting or detecting reasoning errors. These methods overlook the predictive value of explanations. We introduce Normalized Simulatability Gain (NSG), a general and scalable metric based on the idea that a faithful explanation should allow an observer to learn a model's decision-making criteria, and thus better predict its behavior on related inputs. We evaluate 18 frontier proprietary and open-weight models, e.g., Gemini 3, GPT-5.2, and Claude 4.5, on
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית