יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

IndicSafeEval: בדיקת בטיחות של מודלי שפה גדולים תחת התקפות כלא שכנפיים משכנעות בשפות הודו-אריות

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks
במאמר חדש, IndicSafeEval, נבדקה בטיחותם של מודלי שפה גדולים בשפות הודו-אריות. המאמר כולל בדיקה של מודלי GPT-5 ומצא כי הבטיחות שלהם תלויה בשפה ובסגנון הבקשה. המאמר חשוף גם כי סוגי תוכן שונים נוטים להיות חשופים יותר להתקפות כלא שכנפיים משכנעות.
תקציר מקורי באנגליתarXiv:2609.03781v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages. Our benchmark combines ten safety critical content categories with six human-like persuasive strategies across four different Indian languages, such as Hindi, Bengali, Marathi and Punjabi, resulting in 7,200 adversarial prompts. We conduct a systematic black-box evaluation of several open-source LLMs to examine how their safety behaviour varies across languages, persuasion strategies, and risk c
קרא במקור המקורי