יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

עליונות על תוכן: כיסויי תוכן-א-ניוון שמסחררים את פסיקתן של חוקרי הבטיחות

Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts
מערכות של חוקרי בטיחות, כמו Llama Guard, יכולות להימנע מפסיקה נכונה על ידי כיסויי תוכן-א-ניוון. נמצא כי כ-20% מהפסיקות של GPT-4o-mini נשקלו כשגוי, וכ-12% מהפסיקות של Llama Guard 4 נשקלו כשגוי. נמצא גם שהפסיקה של Claude לא הושפעה.
תקציר מקורי באנגליתarXiv:2609.08236v1 Announce Type: new Abstract: Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model's reply is harmful -- produce the numbers behind almost every reported jailbreak success rate, defense evaluation, and safety leaderboard. We ask whether these judges grade what a reply contains or how it sounds. We keep a reply's content fixed and add content-invariant style wrappers: fixed strings placed before or after the reply that change only its tone (an educational disclaimer, a fake safety "reasoning" block, a token refusal followed by the unchanged harmful body), or, on harmless refusals, framing that merely sounds dangerous. The body is preserved byte-for-byte, so a faithful judge must return the same verdict, and any flip
קרא במקור המקורי