יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

FigAct: תיאור דינמי של תמונות מדעיות

FigAct: Turning Scientific Figures into Active Canvases for Explanation
FigAct משנה תמונות מדעיות סטטיות להצגות ויזואליות שתלויות בשאלות. הפרויקט מציע תיאור דינמי של תמונות מדעיות, כמו נואל שמציג תמונות. FigAct מגדיל את היכולת של LLMs להסביר תמונות מדעיות.
תקציר מקורי באנגליתarXiv:2609.36190v1 Announce Type: new Abstract: Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing to
קרא במקור המקורי