יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

סמנטיות: תאוצה והתנגשות גרמים: חישוב בטחון דריפט בייצוגים מערבלים

Semantic Fibers and Cross-Gram Interference: A Calculus of Safety Drift in Overcomplete Representations
במאמר זה נחקרה תופעת דריפט בבטחון דגימות שפה, כאשר דגימות שפה אחת עשויה להיות בטוחה, ודגימות שפה אחרת, שהיא תרגום נאמן, עשויה להיות מסוכנת. המאמר חוקר את הסיבות לדריפט זה ומציע פתרונות לבעיית הבטחון.
תקציר מקורי באנגליתarXiv:2609.14861v1 Announce Type: new Abstract: A deployed language model may refuse a harmful request in English yet comply with its faithful translation, revealing a cross-lingual safety failure that cannot be characterized reliably by output behavior alone. We formalize this phenomenon through an audited equivalence relation and show that, for a declared quotient, representation, metric, feature dictionary, scoring head, threshold, and contrast model, the resulting safety drift admits an exact linear-algebraic characterization. Specifically, the drift is a cross-Gram functional of the within-fiber contrast; its worst admissible value is a support function, while margin invariance is characterized by an annihilator condition. We introduce an intrinsic calibrated exposure measure, governe
קרא במקור המקורי