כתבה
arXiv cs.CL ·
מעבר לצורות פנימיות: טקסונומיה מכניסטית-מקיפה של קודים לשוניים עקיפים לזיהוי שפה קודוו
Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection
טקסונומיה מקיפה לקודים לשוניים עקיפים ב-LLMs. המחקר פיתח טקסונומיה שמסווגת את המנגנונים העומדים בבסיס קודים לשוניים עקיפים, ובדקה את יעילותה בזיהוי שפה קודוו. הטקסונומיה המקיפה והמכניסטית-מקיפה שפותחה במחקר זה, חיסלה את הטקסונומיות הקיימות והציגה תוצאות טובות יותר בזיהוי שפה קודוו.
תקציר מקורי באנגליתarXiv:2606.27314v2 Announce Type: replace Abstract: To avoid moderation and surveillance on social media, some users routinely invent indirect linguistic expressions (ILE) that camouflage sensitive meanings. Such disguised expressions surface as algospeak, euphemisms, and adversarial obfuscation, depending on intent and context, and often involve recurring encoding mechanisms. We propose a comprehensive, mechanism-oriented taxonomy of ILE that abstracts away from communicative goals and instead categorizes the underlying operations through which meaning is encoded and recovered. We evaluate the taxonomy by incorporating it into large language model (LLM) prompts and comparing it with four existing taxonomies and a no-taxonomy baseline, using 2,000 manually annotated TikTok and Bluesky post
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית