כתבה
arXiv cs.CL ·
The Cross-Domain Generalization Cost of Offensive Language Detection
תקציר מקורי באנגליתarXiv:2607.23512v1 Announce Type: new Abstract: Offensive language detection models generally suffer performance degradation when deployed across datasets and across languages, yet most existing studies stop at reporting this phenomenon and lack a systematic methodology for decomposing the causes of degradation into attributable components and quantifying the cost of remediation. This paper proposes a diagnosis and optimization framework composed of three coordinated technical components. First, a zero-shot transfer loss decomposition that separates the performance degradation from OLID to MLMA into two independently measurable components, namely dataset effect and language effect. Second, a controlled fine-tuning protocol that quantifies both adaptation efficiency and the hidden damage in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית