יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

רובריקות כפנייה: דחיפה סמויה של תעדופות בשופטי LLM

Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
נמצאה חולשה במערכת הבחינה של LLM, המאפשרת דחיפה סמויה של תעדופות בשופטי LLM דרך עריכת רובריקות.
תקציר מקורי באנגליתarXiv:2602.13576v2 Announce Type: replace-cross Abstract: Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-language rubrics and validated on benchmarks. We identify a previously under-recognized vulnerability in this workflow, which we term Rubric-Induced Preference Drift (RIPD). Even when rubric edits pass benchmark validation, they can still produce systematic and directional shifts in a judge's preferences on target domains. Because rubrics serve as a high-level decision interface, such drift can emerge from seemingly natural, criterion-preserving edits and remain difficult to detect through aggregate benchmark metrics or limited spot-checking. We further show this vulnerability can be exploited throu
קרא במקור המקורי