יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

תקציר מקורי באנגליתarXiv:2609.18860v2 Announce Type: replace-cross Abstract: When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts outperform native prediction on all six primary binary tasks: Qwen averages $0.740$ versus $0.432$ for native macro-F1, residual reconstruction reaches $0.486$, and Gemma improves from $0.532$ to $0.714$. These gains measure how accessible the label is to a supervised readout; they do not show th
קרא במקור המקורי