יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פרוטוטיפ נייטרלי אינו הנחיה לבטיחות: תלות במקורות והטיות תרגומים באימבדינג של בטיחות-תגובה

A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings
לא ניתן לסקור את בטיחות התגובה על ידי דומיננטיות קוסינוס לאימבדינג-ממוצע של תגובות בטוחות. המאמר חוקר את הטענה על שני קבוצות-פיקד ואחד קבוצת-משפט, באמצעות ארבעה אימבדרים-קפואים ופיצולי-תרגום. התוצאות מצביעות על חשיפה של נקודת-משיכה-חיובית-בלתי-מוגדרת, וכי נקודת-משיכה-חיובית-מוגדרת-ברורה יכולה לספור בטיחות תגובה טוב יותר.
תקציר מקורי באנגליתarXiv:2610.01801v1 Announce Type: new Abstract: Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings
קרא במקור המקורי