יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מאפיינים דלים: חסרון קריאה במודלי תצוגה-שפה לזיהוי ממים פוגעניים

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
מודלי תצוגה-שפה קשים לזיהוי ממים פוגעניים, וזאת עקב חסרון אווידנטים פנימיים או בעיות רוטינג. החידוש: ניתוח עם מאפיינים דלים, פרובים מותנים, והתערבויות סיבתיות. המודלים: Gemma-3 ו-Qwen3.5.
תקציר מקורי באנגליתarXiv:2609.18860v2 Announce Type: replace-cross Abstract: When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts outperform native prediction on all six primary binary tasks: Qwen averages $0.740$ versus $0.432$ for native macro-F1, residual reconstruction reaches $0.486$, and Gemma improves from $0.532$ to $0.714$. These gains measure how accessible the label is to a supervised readout; they do not show th
קרא במקור המקורי