כתבה
arXiv cs.LG ·
מודלי שמירה עדיפים כאשר מודלי בסיס הם ספקנים
Guard Models Are Overconfident Where Base Models Are Uncertain
מודלי שמירה עדיפים נוטים להיות עדיפים כאשר מודלי בסיס הם ספקנים. זה נמצא במאמר שפורסם ב-arXiv:2609.36477v1.
תקציר מקורי באנגליתarXiv:2609.36477v1 Announce Type: new Abstract: Guard models are used as safety classifiers, with confidence scores driving downstream moderation decisions. We evaluate five guard models for prompt classification and find that although several are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections. Comparing each guard with its corresponding base LM, we find that uncertainty signals often remain available, with the base model typically expressing uncertainty on the same inputs where the guard fails. Layer-wise analyses localize this guard-base divergence to later layers, where guard models exhibit sharper safe/unsafe separation and lower-rank
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית