יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הכשל הוא בקריאה: מדדי הזיהוי המדויק של הרגשות נמדדים, ולא תפיסה

The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception
מדדי הזיהוי המדויק של הרגשות מציגים את VLMs כבורות בביצועיהם כאשר מבקשים מהם לספק תשובות, אך טובים כאשר קוראים מהלוגיטים.
תקציר מקורי באנגליתarXiv:2610.08162v1 Announce Type: cross Abstract: Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark's images, taxonomy and ratings, and change only how
קרא במקור המקורי