כתבה
arXiv cs.AI ·
Why Do Vision Language Models Struggle To Recognize Human Emotions?
תקציר מקורי באנגליתarXiv:2604.15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) have made tremendous progress in the last few years for many visual tasks, potentially offering a promising solution for understanding emotions. However, it is surprising that even the most sophisticated contemporary VLMs struggle to recognize human emotions or to outperform even specialized vision-only classifiers. In this paper we ask the question 'Why do VLMs struggle to recognize human emotions?', and observe that the inherently continuous and dynamic task of facial expression recognition (DFER) exposes two critical VLM vulnerabilities. First, emotion datasets are naturally long-tailed, and th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית