כתבה
arXiv cs.CL ·
VisionFoundry: הוראת תפיסה חזותית ל-VLMs
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
VisionFoundry היא פלטפורמה אוטומטית להוראת תפיסה חזותית למודלים של שפה וחזון (VLMs) באמצעות תמונות סינתטיות. היא משתמשת במודלים של שפה (LLMs) כדי ליצור שאילתות ותמונות סינתטיות, ומשפרת את היכולות החזותיות של VLMs כמו Qwen.
תקציר מקורי באנגליתarXiv:2604.09531v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition, largely because natural image datasets provide limited supervision for low-level visual skills. Can targeted synthetic supervision address these weaknesses without reference images or manual annotation? To investigate this, we introduce VisionFoundry, an automated pipeline that takes only a task name as input, uses LLMs to synthesize paired questions, answers, and text-to-image (T2I) prompts, generates images with T2I models, and filters samples via multimodal verification. With VisionFoundry, we construct VisionFoundry-10k, a synthetic VQA dataset spanning 10 perception tasks. Finetuning on VisionFoundr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית