כתבה
arXiv cs.AI ·
VisionQ: מערך ומבחן לביקורת חזותית מורכבת
VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision
VisionQ הוא מערך ומבחן חדש לביקורת חזותית מורכבת, המאפשרת למודלי VLM להגדיר קריטריונים חזותיים ולבחון את התוצאות. המערך כולל 1,409 מאמרים ו-1,800+ תמונות ומספק תוצאות דיוק של 63.1%.
תקציר מקורי באנגליתarXiv:2610.00666v1 Announce Type: cross Abstract: Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visual reason. We introduce VisionQ, the first benchmark built from peer-reviewed CV comparison figures that grounds every judgment in a named visual criterion: each question states the criterion, and a judge is credited only when it selects the output the authors identify as best on that criterion. We call this task criterion-conditioned visual discrimination. VisionQ comprises (1) a corpus of 1,409 CVPR and ICCV papers with 1,800+ validated compari
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית