כתבה
arXiv cs.CL ·
Chehre: מאגר נתונים לחקר רגישות תצפיתית במודלי שפה וידאו
Chehre: An Emoji-Prompted Dataset to Explore Perceptual Flexibility in Video Language Models
מאגר נתונים חדש לחקר רגישות תצפיתית במודלי שפה וידאו. המאגר כולל סרטוני וידאו של 203 נבדקים שהציגו 40 פנים שונות. הסרטונים תורגמו לפנים סינתטיות כדי לשמור על פרטיות. 1,242 נבדקים תיארו את הסרטונים, והתוצאות נמצאו ל-2,111 סרטונים.
תקציר מקורי באנגליתarXiv:2606.21657v2 Announce Type: replace-cross Abstract: Do people perceive the same facial expression in the same way? Should we expect vision models to be flexible in how they perceive facial expressions? Facial expressions are nonverbal social signals used in human interaction, but facial expression recognition datasets often focus on a single deterministic annotation per sample. We introduce Chehre, an emoji-prompted video dataset with a wide range of dynamic facial expressions for exploring perceptual variation. In Chehre, 203 participants were prompted to express and record 40 facial emojis. Later, their facial motions were transferred onto synthetic faces to preserve privacy. A separate group annotated the videos, resulting in 2,111 videos annotated by 1,242 perceivers, with ~30 an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית