כתבה
arXiv cs.CL ·
מדידה ושיפור ההתאמה בין CoT לפענוח LLM
Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment
חוקרים מציגים שיטה למדידה ושיפור ההתאמה בין Chain-of-thought (CoT) לפענוח LLM. המחקר בודק את ההתאמה בין CoT לפענוח פנימי במודלים שונים, כולל LLMs. התוצאות מראות כי ניתן לשפר משמעותית את ההתאמה בין CoT לפענוח פנימי.
תקציר מקורי באנגליתarXiv:2609.38972v1 Announce Type: new Abstract: Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure and improve the alignment between the reasoning described in an LLM's CoT and what it computes internally. We propose CoT-Interpretability Alignment (CIA), a metric that measures the agreement between a model's CoT traces and its internal reasoning strategies as detected by interpretability tools. We evaluate CIA on three tasks (two-hop question answering, hint intervention, and integer multiplication) across three LLMs, finding that LLMs exhibit lim
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית