כתבה
arXiv cs.AI ·
Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers
תקציר מקורי באנגליתarXiv:2603.01437v2 Announce Type: replace Abstract: As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, suggesting the opportunity to understand model decisions through verbalized reasoning. However, the utility of CoT toward interpretability depends upon its faithfulness---whether the model's stated reasoning reflects the underlying decision process. We provide mechanistic evidence that instruction-tuned models often determine their answer before generating CoT. Training linear probes on residual stream activations at the last token before CoT, we can predict the model's final answer with >0.9 AUC on most tasks. We find that these directions are not only predictive, bu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית