כתבה
arXiv cs.AI ·
חקירת פרומפט רב-מודאלי ליצירת ויזואליזציה
Exploring Multimodal Prompt for Visualization Authoring with Large Language Models
חוקרים את הפוטנציאל של פרומפטים רב-מודאליים ליצירת ויזואליזציה. המחקר מציג את VisPilot, כלי שמאפשר למשתמשים ליצור ויזואליזציות באמצעות פרומפטים רב-מודאליים, כולל טקסט, רישומים ומניפולציות ישירות על ויזואליזציות קיימות. התוצאות מראות כי פרומפטים רב-מודאליים מסייעים למשתמשים לתקשר את הכוונה הוויזואלית בצורה יעילה יותר.
תקציר מקורי באנגליתarXiv:2504.13700v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language is limited in precision and expressiveness for conveying visualization intent, leading to misinterpretation and time-consuming iterations. To address these limitations, we conduct an empirical study to understand how LLMs interpret ambiguous or incomplete text prompts in the context of visualization authoring, and the conditions making LLMs misinterpret user intent. Informed by the findings, we introduce visual prompts as a complementary input modality to text prompts, which help clarify user intent and improve LLMs
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית