כתבה
arXiv cs.CL ·
התקן Percept-V: האם LLMs המודעים-מודליים יכולים לפצח בעיות תפיסה פשוטות?
The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
במאמר זה, נבחן את יכולתם של LLMs המודעים-מודליים לפצח בעיות תפיסה פשוטות. התקן Percept-V, המכיל 6,000 תמונות תוכנתיות, נועד לבדוק את יכולתם של LLMs להבין ולהפוך תמונות. התוצאות היו חסרוניות, וה-LLMs לא הצליחו להפוך את התמונות באופן טוב.
תקציר מקורי באנגליתarXiv:2508.21143v4 Announce Type: replace Abstract: Cognitive science research treats visual perception, the ability to understand and make sense of a visual input, as one of the early developmental signs of intelligence. Its TVPS-4 framework categorizes and tests human perception into seven skills such as visual discrimination, and form constancy. Do Multimodal Large Language Models (MLLMs) match up to humans in basic perception? Even though many benchmarks evaluate MLLMs on advanced reasoning and knowledge skills, there is limited research that focuses evaluation on simple perception. In response, we introduce Percept-V, a dataset containing 6000 program-generated uncontaminated images divided into 30 domains, where each domain tests one or more TVPS-4 skills. Our focus is on perception,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית