כתבה
arXiv cs.AI ·
AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths
תקציר מקורי באנגליתarXiv:2609.06771v1 Announce Type: cross Abstract: Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism analysis, account linking, misinformation investigation, and machine-generated text detection. Yet current authorship benchmarks remain fragmented, usually covering only a narrow language set, a single genre, or a limited document-length regime, which makes it difficult to assess whether modern representations truly generalize. We introduce AuthBench, a large-scale multilingual benchmark for authorship representation that is designed to make this evaluation broad, standardized, and realistic. AuthBench contains 428,150 documents written by 153,825 individuals across ten widely used languages, 9 primary genres, 66 fine-grained genres, an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית