יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SAE++: קסדה של אוטואנקודרים דקות לומר רמות רבות של רעיונות נוף ב-MLLMs

SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs
קסדה של אוטואנקודרים דקות לומר רמות רבות של רעיונות נוף ב-MLLMs. המחקר עוסק בפיתוח קסדה של אוטואנקודרים דקות שמסוגלים ללמוד רעיונות נוף רב-רמתיים ב-MLLMs. הקסדה נבנית על ידי רשת עם שני רמות: רמה ראשונה של אוטואנקודרים דקות ורמה שנייה של אוטואנקודרים דקות שמסוגלים ללמוד רעיונות נוף רב-רמתיים. הקסדה נבחנה במספר נתוני ניסויים והתוצאות הראו כי הקסדה מסוגלת ללמוד רעיונות נוף רב-רמתיים ב-MLLMs.
תקציר מקורי באנגליתarXiv:2606.16193v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret. Sparse Autoencoders (SAEs) provide a scalable way to decompose dense model activations into sparse, interpretable features. However, existing SAE architectures primarily recover flat feature dictionaries and are less suited for explicit multi-level concept organization. In this paper, we introduce a cascaded sparse autoencoder architecture, dubbed SAE++, for learning hierarchical visual concepts in MLLMs. Rather than nesting or stacking SAE sparse activation codes, SAE++ trains a second-level SAE directly on the decoder weights of the first-level SAE, treatin
קרא במקור המקורי