כתבה
arXiv cs.AI ·
VisionFoundry: חינוך דגמי VLM לראייה עם תמונות סינתטיות
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
VisionFoundry: חינוך דגמי VLM לראייה עם תמונות סינתטיות. פיתוח פילוסופיה חדשה לאימון VLMs עם תמונות סינתטיות.
תקציר מקורי באנגליתarXiv:2604.09531v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition, largely because natural image datasets provide limited supervision for low-level visual skills. Can targeted synthetic supervision address these weaknesses without reference images or manual annotation? To investigate this, we introduce VisionFoundry, an automated pipeline that takes only a task name as input, uses LLMs to synthesize paired questions, answers, and text-to-image (T2I) prompts, generates images with T2I models, and filters samples via multimodal verification. With VisionFoundry, we construct VisionFoundry-10k, a synthetic VQA dataset spanning 10 perception tasks. Finetuning on VisionFoundr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית