כתבה
arXiv cs.LG ·
Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
תקציר מקורי באנגליתarXiv:2610.10625v1 Announce Type: cross Abstract: Looped Language Models (LoopLMs) provide a parameter efficient approach to scaling model capabilities through repeated use of shared parameters across recurrent steps. Since each recurrent depth can be read out independently, a single LoopLM exposes a broader output space across inference depths, raising an important question: whether safety is preserved throughout recurrent computation. Prior evaluations suggest that deeper recurrence can improve safety on harmful queries, but robustness under jailbreak attacks remains unclear. We therefore conduct a comprehensive safety evaluation of LoopLMs under jailbreak attacks targeting different recurrent depths. We find that attack success can increase at deeper inference depths, the same query can
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית