יום חמישי, 30 ביולי 2026 LIVE
AI־INFO

כתבה MarkTechPost ·

Zyphra משחררת Zamba2-VL: מודלים היברידיים לראייה-שפה

Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order of Magnitude
Zyphra השחררה את Zamba2-VL, מודל ראייה-שפה היברידי המשלב יכולות ראייה ושפה. המודל מבוסס על ארכיטקטורת Zamba2 היברידית, הכוללת שכבות Mamba2 ובלוקים משותפים. Zamba2-VL תומכת בהבנה והקשר של תמונות וטקסט.
תקציר מקורי באנגליתZyphra has released Zamba2-VL, a family of open vision-language models. The release covers three sizes: 1.2B, 2.7B, and 7B parameters. Each model is built on the Zamba2 hybrid SSM–Transformer backbone. Vision-language models (VLMs) read images and text together. They answer questions about charts, documents, and photos. Most open VLMs use a dense Transformer as the language model. Zamba2-VL replaces that with a hybrid state-space design. The goal is competitive accuracy at lower latency. What is Zamba2-VL Zamba2-VL follows the now-standard LLaVA-style VLM template. A pre-trained vision encoder turns image patches into features. A lightweight MLP adapter projects those features into the language model’s space. The language model then reads an interleaved sequence of vision and text to
קרא במקור המקורי