כתבה
arXiv cs.AI ·
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
תקציר מקורי באנגליתarXiv:2607.25527v1 Announce Type: cross Abstract: Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities. To address these challenges, we present Argus-Unified, a compact, effective and unified multimodal model built with low demand on computation and data. Instead of aligning modalities from scratch, Argus-Unified effectively leverages pretrained vision-language models (VLMs) that provide strong multimodal priors. Specifically, we introduce hybrid visual tokens that preserve continuous tokens for understanding while learning discrete tokens for generation from a frozen unified vision encoder. Our training pipeline
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית