כתבה
arXiv cs.AI ·
IndicSafeEval: בדיקת אמינות של מודלי שפה גדולים תחת התקפות כלאבקה משכנעות בשפות הודו-אירופיות
IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks
בדיקת אמינות של מודלי שפה גדולים תחת התקפות כלאבקה משכנעות בשפות הודו-אירופיות. המחקר חושף חולשות בבדיקות אמינות קיימות, המתמקדות בעיקר באנגלית.
תקציר מקורי באנגליתarXiv:2609.03781v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages. Our benchmark combines ten safety critical content categories with six human-like persuasive strategies across four different Indian languages, such as Hindi, Bengali, Marathi and Punjabi, resulting in 7,200 adversarial prompts. We conduct a systematic black-box evaluation of several open-source LLMs to examine how their safety behaviour varies across languages, persuasion strategies, and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית