יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

VisionWeave: ייצוגים חזותיים גמישים

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
VisionWeave היא טכנולוגיה חדשה המאפשרת ייצוגים חזותיים גמישים במודלים רב-מודאליים. היא משלבת שני מרכיבים: גייטד ספיישל פולר וראוטר עדינות. היא פותחה על ידי Qwen3.5-4B ו-Qwen3.8-27B.
תקציר מקורי באנגליתarXiv:2610.07987v1 Announce Type: cross Abstract: Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving.
קרא במקור המקורי