כתבה
arXiv cs.AI ·
יצירת SVG מורכבת דרך פרסינג סמנטי היררכי
Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing
חוקרים מציגים פריימוורק ליצירת גרפיקה וקטורית מורכבת באמצעות מודלים ויז'ואלי-לשוניים. הפריימוורק מאפשר יצירת SVG מורכבים ועריכה קלה יותר. החוקרים גם מציגים בנך' מרק להערכת איכות היצירה.
תקציר מקורי באנגליתarXiv:2609.14657v1 Announce Type: cross Abstract: While Vision-Language Models (VLMs) excel at visual reasoning, generating structured, editable Scalable Vector Graphics (SVG) remains a fundamental challenge. Existing pipelines predominantly yield flat, semantically agnostic collections of paths, where editing a single object requires manually identifying its constituent paths. To address this, we propose a VLM-driven agentic framework for semantic compositional SVG generation. Our pipeline recursively parses visual scenes into semantic and geometric hierarchies via top-down decomposition, visual grounding, and prompt-driven amodal occlusion recovery, ensuring each component is geometrically complete. Furthermore, we introduce the Semantic SVG Benchmark with human-annotated semantic groups
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית