יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בדיקת בטיחות של מודלי שפה גדולים עם פניות לא-קנוניות

Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs
בדיקת בטיחות של מודלי שפה גדולים עם פניות לא-קנוניות. המאמר עוסק בבדיקת בטיחות של מודלי שפה גדולים עם פניות לא-קנוניות. המחברים פיתחו דטסט חדש של 2,100 פניות ובדקו 5 מודלי שפה גדולים. התוצאות הראו שהמודלים יכולים להפיק תשובות הרתעותיות כאשר הפניות כוללות סימניות, כתיב חריג, ועוד.
תקציר מקורי באנגליתarXiv:2610.09033v1 Announce Type: new Abstract: Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with pool
קרא במקור המקורי