כתבה
arXiv cs.CL ·
לבטל תכונות גזעניות, לפגוע ללא סיבה: התקפות סגנון-מסמן-על-LLM-בעבור-מודרציה
Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation
מחקר מצא ש-LLMs הם רגישים להתקפות סגנון-מסמן-על שמעבירות את היכולת למודרציה. התקפות אלה יכולות לפגוע ב-LLMs, ולכן יש צורך בביטחון ובהערכה נפרדים בעבור-מודרציה.
תקציר מקורי באנגליתarXiv:2608.22230v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such feedback introduces two manipulation directions: whitewashing hateful content as normal and smearing normal content as hateful. This study examines the susceptibility of initially correct model judgments to annotator-style rebuttals and analyzes whether attack effectiveness differs across manipulation directions. We introduce a rejudge protocol that extends direct contradiction with decision-boundary perturbations and adversarial rationales. Experiments with multiple LLMs on two hate speech datasets show that annotator-style rebuttals substantially degrade moderat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית