כתבה
arXiv cs.AI ·
On the Chain-of-Thought Monitorability of Looped Language Models
תקציר מקורי באנגליתarXiv:2610.02741v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring provides a promising approach for detecting undesirable model behavior. Looped language models (LoopLMs) repeatedly apply shared transformer layers, increasing effective computational depth and enabling additional latent computation without increasing model size. However, the effect of looped architectures on CoT monitorability remains largely unexplored. In this work, we provide the first systematic evaluation of CoT monitorability in LoopLMs. We study two complementary settings: (1) varying the loop depth within the same LoopLM family to isolate the effect of additional recurrent computation, and (2) comparing LoopLMs with non-looped language models matched by parameter size, transformer-layer count, or eff
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית