כתבה
Simon Willison ·
צוות החזית של Anthropic
Quoting Anthropic Frontier Red Team
צוות החזית של Anthropic בדק מודלים GLM-5.3 ו-Claude Mythos Preview ב-100 משימות ניצול בינארי. GLM-5.3 הצליחה ב-4% מהניסיונות, בעוד Claude Mythos Preview הצליחה ב-6%. התוצאות מראות על חציית סף משמעותי ביכולת המודלים לבצע התקפות סייבר מתקדמות.
תקציר מקורי באנגלית<blockquote cite="https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities"><p>We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.</p></blockquote> <p class="cite">— <a href="https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities">Anthropic Frontier Red Team</a>, GLM-5.3 and the spread of advanced cyber capabilities</p> <p>Tags: <a href="https://simonwillison.ne
קרא במקור המקורי
simonwillison.net
פתח כתבה מקורית