כתבה
arXiv cs.CL ·
לא כל היגיון של LLM נראה בשרשרת המחשבה
Not All LLM Reasoning is Visible in the Chain-of-Thought
חוקרים גילו שמודלים מתקדמים כמו Claude יכולים לבצע חישובים משמעותיים ללא עקבות מפורשת בפלט. התופעה נצפתה במטלות סינתטיות ומראה על אתגרים בנושא בטיחות AI.
תקציר מקורי באנגליתarXiv:2607.22925v2 Announce Type: replace Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on synthetic reasoning tasks. We evaluate 13 frontier language models across three tasks and find that many models benefit significantly from filler tokens, with accuracy improvements of up to 13 percentage points. The benefit depends on which tokens are used and differs across models. We further show that filler tokens enable Claude Opus 4.5 to satisfy a hidden modular arithmetic constraint without sacrificing accuracy on its primary task, demonstrating that invisible
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית