כתבה
arXiv cs.AI ·
SalamahBench: Dialect and Category Level Safety Evaluation of Arabic Language Models
תקציר מקורי באנגליתarXiv:2603.04410v3 Announce Type: replace-cross Abstract: While different stakeholders are trying to leverage Arabic Language Models (ALMs), safety alignment in ALMs remains largely underexplored, hindering their mainstream adoption. Existing safety benchmarks are predominantly English-centric and evaluate Arabic only in its standardized form, obscuring fine-grained safety vulnerabilities in Arabic NLP systems. This paper introduces SalamahBench, a unified benchmark of 8{,}270 human-verified harmful prompts across ML Commons hazard categories, each rendered in Modern Standard Arabic (MSA) and five regional Arabic varieties, namely Egyptian, Syrian, Saudi, Lebanese, and Moroccan, for a total of 49{,}620 paired instances. To analyze the resulting data, we introduce two complementary metrics,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית