כתבה
arXiv cs.CL ·
Strong Multilingual Privacy Tagging at Encoder Speed
תקציר מקורי באנגליתarXiv:2609.38630v1 Announce Type: new Abstract: Privacy redaction must remove personal information while preserving relationships expressed in text. We develop a multilingual named-entity tagger with fine-grained distinctions supporting varied redaction policies and methods for cheaply learning additional distinctions. We fine-tune a multilingual encoder with an affine span-tagging head on frontier-model annotations in 35 languages, replay mapped human gold with coverage-aware masking so unannotated types are not treated as negatives, and repair subword boundaries with a learned +/-1-character adjustment. On 1,283 human-gold test segments in seven languages, best measured redaction F1 is 88.8, against 69.1 for published GLiNER2 with 11 unrepresentable types excluded from its task (68.8 wit
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית