יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Transcoders: חידוש בהבנת תהליך התרגום במודלי תצלום-לשפה

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models
מחקר חדש חושף את התהליך המדויק שבו מודלי תצלום-לשפה תרגמים תמונות לטקסט. המחקר משתמש בטכניקה חדשה של transcoders, שמאפשרת הבנה טובה יותר של התהליך המדויק שבו המודלים תרגמים תמונות לטקסט. המחקר גם מציג תוצאות של ניתוח סטרוקטורלי של התוצאות של המודלים, ומציע פתרונות להבנת התהליך המדויק שבו המודלים תרגמים תמונות לטקסט.
תקציר מקורי באנגליתarXiv:2605.22902v2 Announce Type: replace-cross Abstract: Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood. Existing interpretability work on VLMs uses Sparse Autoencoders (SAEs), which decompose static residual representations and miss the functional updates that drive cross-modal interaction. We adopt a function-centric framework based on Transcoders, sparse approximations of MLP sublayers that act as a causal proxy for layer-wise computation. Applied to Gemma 3-4B-IT, the framework decomposes the model into interpretable computational pathways linking image patches to directions in token generation. Transcoder attributions produce stronger and more stable effects on visually grounded to
קרא במקור המקורי