יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

רוביקס כפנייה: נטיית תערובת בשופטי LLM

Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
נטיית תערובת סתרית בשופטי LLM דרך עריכות רוביקס. ניתן להשתמש בעריכות רוביקס שאינן פוגעות בבסיס המבחן כדי לשנות את העדפות השופט. זה עלול לגרום לדעיכה בדיוק של 9.5% (עדיפות) ו-27.9% (לא-עדיפות) בתחום המטרה.
תקציר מקורי באנגליתarXiv:2602.13576v2 Announce Type: replace-cross Abstract: Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-language rubrics and validated on benchmarks. We identify a previously under-recognized vulnerability in this workflow, which we term Rubric-Induced Preference Drift (RIPD). Even when rubric edits pass benchmark validation, they can still produce systematic and directional shifts in a judge's preferences on target domains. Because rubrics serve as a high-level decision interface, such drift can emerge from seemingly natural, criterion-preserving edits and remain difficult to detect through aggregate benchmark metrics or limited spot-checking. We further show this vulnerability can be exploited throu
קרא במקור המקורי