יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בדיקת בטיחות רב-מצבית

Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs
המחקר מציג את ASRD, מאגר נתונים לבדיקת בטיחות של מודלים גדולים. הוא בודק חמישה מודלים, כולל Mistral 7B, על 2,100 קלטים שונים. התוצאות מראות כי המודלים מגיבים שונה על קלטים לא-קנוניים.
תקציר מקורי באנגליתarXiv:2610.09033v1 Announce Type: cross Abstract: Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with po
קרא במקור המקורי