כתבה
arXiv cs.CL ·
SWARM: קבוצת נתונים רב-לשונית לזיהוי תעמולה רוסית בתוצאות מנועי חיפוש
SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results
קבוצת נתונים לזיהוי תעמולה רוסית בתוצאות מנועי חיפוש. הנתונים כוללים 2,183 תוצאות מנועי חיפוש בתשע שפות ובאתרי אינטרנט שונים. הקבוצת הנתונים תוכל לסייע בזיהוי תעמולה רוסית בתוצאות מנועי חיפוש.
תקציר מקורי באנגליתarXiv:2609.12653v1 Announce Type: new Abstract: Russian state propaganda spreads across many languages and online spaces. Yet, most computational work examines only one such space, usually social media, in one or two languages, and analyses sources rather than content. We introduce SWARM (Search-Web documents Annotated for Russian propaganda, Multilingual), a dataset of 2,183 search engine results across nine languages and diverse web domains (e.g., news, blogs, government sites), each annotated by trained coders for whether it supports a recurring Russian propaganda narrative. We benchmark a source-based blocklist, supervised classifiers, and zero-shot LLMs against these labels. The blocklist misses most propaganda-supporting documents, because such content is not confined to flagged "pro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית