יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

התקפות מפאת צופן שרירותיות נגד מודלי שפה גדולים

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning
חוקרים הדגימו התקפות מפאת צופן שרירותיות נגד מודלי שפה גדולים של Anthropic, Google ו-OpenAI, ללא צורך בקליברציה מחדש. ההתקפות מאפשרות תקשורת סמויה דרך המודלים, ומערערות את האבטחה שלהם.
תקציר מקורי באנגליתarXiv:2609.09553v1 Announce Type: cross Abstract: Large language model safety and security research is preoccupied with, among other things, detecting and preventing jailbreak attacks: alignment bypasses that allow an adversarial user to elicit unwanted or harmful outputs from models. Arbitrary cipher, or covert communication, attacks are one such type of jailbreak and have previously been demonstrated against the fine-tuning APIs of commercial models. In these attacks, target models are trained on a corpus of encrypted harmful questions and responses and subsequently respond to harmful requests through the learned encryption scheme. In this paper, we show that newer frontier models do not require fine-tuning to acquire cipher-based communication skills. Instead, they can learn these skill
קרא במקור המקורי