יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

NeuronEye: פענוח ויזואלי-שפתי של קונצפטים ויזואליים

NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning
NeuronEye היא פלטפורמה שמאפשרת פענוח ויזואלי-שפתי של קונצפטים ויזואליים. היא משתמשת בשאילתות שפה כדי לפענח ולמקד קונצפטים ויזואליים בתמונות.
תקציר מקורי באנגליתarXiv:2609.38098v1 Announce Type: new Abstract: Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language que
קרא במקור המקורי