כתבה
arXiv cs.LG ·
Do MLLM Judges Judge the Edit? - ניקיון שיפוטי של יורים MLLM
Do MLLM Judges Judge the Edit? Auditing Bias in Image Editing Evaluation with Verified Quality Preservation
אודיט שיפוטי של MLLM לביקורת על עריכת תמונות עם אימות של שימור תכונות יצירה
תקציר מקורי באנגליתarXiv:2610.01670v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model training. However, systematically auditing whether these judges are influenced by cues irrelevant to editing quality is challenging because visual interventions may themselves alter the quality being evaluated. A judgment shift can therefore be attributed to bias only when the intervention is verified to preserve the underlying editing quality. To address this challenge, we introduce EditJudgeBias, a counterfactual benchmark with verified quality preservation, comprising 1,196 real editing samples and 13 cues injected across four evaluation sites. We verify quality preservation for the reques
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית