יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חיפוש אחר "הסירוב המזיק": סקירה פסיכומטרית של מבחן בטיחות ל-AI

Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark
בדיקה פסיכומטרית של מבחן בטיחות ל-AI מגלה ששיטות הנוכחיות עשויות להתערבב בין התנהגויות מזיקות.
תקציר מקורי באנגליתarXiv:2610.12409v1 Announce Type: new Abstract: Safety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profiles. Comparing models is more tractable at the level of individual attributes, yet it is often unclear whether even a single dataset's scores isolate any single attribute. One plausible candidate for such an attribute is harmful refusal, a model's tendency to refuse dangerous or policy-violating prompts. We examine whether it constitutes a single, measurable attribute in HELM Safety. Using a construct validity framework that stipulates that an attribute must exist before a test can measure it, we start with HELM Safety's four d
קרא במקור המקורי