כתבה
arXiv cs.CL ·
חקירה מניפסטית של נוירונים AI-Text ב-BERT: חקירה דקה ותיקון פטצ' של פעילות ב-RAID
A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
במאמר זה, חוקרים חקרו נוירונים ב-BERT המזהים טקסט AI. הם גילו קבוצה קטנה של נוירונים המשמשים לזיהוי טקסט AI, והראו שהם קשורים להתאמה של המודל למשימות AI.
תקציר מקורי באנגליתarXiv:2609.30287v1 Announce Type: new Abstract: AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spanning pure-base and instruction-tuned models. We apply the L1-to-L2 sparse-probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden-state dimensions (12 layers x 768), which we call neurons. The procedure recovers a stable set of under 1% of neurons per generator, consistent across folds and seeds; a probe restricted to that set retains most of the full-feature detection accuracy. Bidirectional activation patching confirms this set's causal
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית