כתבה
arXiv cs.AI ·
Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs
תקציר מקורי באנגליתarXiv:2610.07646v1 Announce Type: new Abstract: Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm from comparative and developmental psychology and pairing it with a mechanistic analysis of the model's internals. On a parametrically controlled stimulus set evaluated across frontier API models (GPT, Claude, Gemini) and three open-source families (Qwen3.5, Gemma-4, InternVL3), we identify four levers that shift VLMs toward the relational match---capability tier, model scale, the number of objects per scene, and the absence of pe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית