כתבה
arXiv cs.LG ·
GroundBench: בנק אבולוציה חדש לבדיקת כשלי תפקוד של ויז'ן-לנגואג'
GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures
בנק אבולוציה חדש לבדיקת כשלי תפקוד של ויז'ן-לנגואג' - GroundBench. הבנק כולל 1,068 תרגילים שבהם נבדקים כשלי תפקוד של ויז'ן-לנגואג' בפני שלושה מודלי OpenAI.
תקציר מקורי באנגליתarXiv:2609.13308v1 Announce Type: cross Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language models, with no model outperforming a constant baseline until the part was named. However, naming the part supplies information that a real system must infer, confounding visual grounding, mechanical reasoning, and category-to-action association. We introduce GroundBench, a diagnostic benchmark that separates these explanations through six branch-and-merge conditions, each adding a controlled information bundle, and a counterfactual re-ask targeting a real alternate part visible in the same image. Across three OpenAI models and 1,068 predictions, supplying the target region without its identit
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית