כתבה
arXiv cs.CL ·
Self-Explaining Hate Speech Detection with Moral Rationales
תקציר מקורי באנגליתarXiv:2601.03481v2 Announce Type: replace Abstract: Existing hate speech detection models are often opaque and rely on surface-level lexical cues, which makes them vulnerable to spurious correlations and limits robustness, interpretability and cultural contextualization. We propose Supervised Moral Rationale Attention (SMRA), the first self-explaining hate speech detection framework to incorporate moral rationales as direct supervision for attention alignment. Based on Moral Foundations Theory, SMRA aligns token-level attention with expert-annotated moral rationales, guiding models to attend to morally salient spans. Unlike prior rationale-supervised or post-hoc approaches, SMRA integrates moral rationale supervision directly into the training objective, producing inherently interpretable
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית