כתבה
arXiv cs.CL ·
הערכת התנהגות מותנית במודלים גדולים של שפה
Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation
חוקרים פיתחו DECO, שיטה להערכת מודלים גדולים של שפה בהתאם לקריטריונים שונים. המחקר מצא שמודלים כמו LLMs ו-GPT מתקשים ליישם קריטריונים ספציפיים.
תקציר מקורי באנגליתarXiv:2609.03814v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply each criterion when making decisions. To study whether LLMs exhibit criterion-conditioned behaviour, we introduce Diagnostic Evaluation of COntent (DECO), a criterion-independent factorisation of content that enables controlled, criterion-level evaluation. We also introduce pairwise evaluation to compare model outputs across different criteria for the same input. Across four moderation datasets and four LLMs, we find that strong benchmark performance can hide substantial failures
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית