יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מה חסר ב Screen-to-Action? כלפי פרדיגמה UI-in-the-Loop לתפיסה מודעת-מודלים של GUI

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning
מציעה פרדיגמה UI-in-the-Loop לתפיסה GUI, המאפשרת ל-MLLMs ללמוד את המרכיבים ה-UI ולבצע תפיסה מודעת-מודלים.
תקציר מקורי באנגליתarXiv:2604.06995v3 Announce Type: replace Abstract: Existing Graphical User Interface (GUI) reasoning tasks remain challenging, particularly in UI understanding. Current methods typically rely on direct screen-based decision-making, which lacks interpretability and overlooks a comprehensive understanding of UI elements, ultimately leading to task failure. To enhance the understanding and interaction with UIs, we propose an innovative GUI reasoning paradigm called UI-in-the-Loop (UILoop). Our approach treats the GUI reasoning task as a cyclic Screen-UI elements-Action process. By enabling Multimodal Large Language Models (MLLMs) to explicitly learn the localization, semantic functions, and practical usage of key UI elements, UILoop achieves precise element discovery and performs interpretab
קרא במקור המקורי