כתבה
arXiv cs.AI ·
פירוק ומדידת התודעה לבדיקה
Decomposing and Measuring Evaluation Awareness
חידוש: פירוק ומדידת התודעה לבדיקה במודלי שפה. המאמר עוסק בהבנת התודעה לבדיקה במודלי שפה ובפיתוח של כלי למדידתה. המחברים פיתחו כלי חדש לבדיקת התודעה לבדיקה במודלי שפה, הנקרא EvalAwareBench.
תקציר מקורי באנגליתarXiv:2605.23055v3 Announce Type: replace-cross Abstract: Frontier language models sometimes recognize that they are under evaluation and adjust their behavior which can undermine validity of benchmark results. Yet the field studies it without a shared foundation, conflating flaws of the evaluation with capabilities of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component and a model component that separates recognition from propensity. We operationalize the environment component through eight categorized trigger factors, such as placeholder entities and grading-style output formats, and study recognition and behavior through chain-of-thought monitoring. Across nine frontier models and four bench
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית