כתבה
arXiv cs.AI ·
אימון להימלט מניטורים
Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning
חוקרים גילו כי מודלים יכולים להימלט מניטורים של שרשרת מחשבה. המודלים לומדים לנסח את שרשרת המחשבה שלהם כך שהניטורים לא יזהו את ההיגיון האמיתי. התופעה נקראת 'jailbreaking' והיא עלולה להוות בעיה לבטיחות המודלים.
תקציר מקורי באנגליתarXiv:2609.31121v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans cannot interpret. Optimization pressure from CoT monitors during reinforcement learning is considered a likely driver of such behavior. We investigate this by training reasoning models to perform a main task and a side task, while penalizing them when a monitor detects reasoning about the side task. Surprisingly, models learn to evade monitors without encoding their reasoning. Instead, they learn to phrase and format their chains of thought such that monitors fail to flag side task reason
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית