כתבה
arXiv cs.AI ·
הגברת GUI Agents עם מעברי מצב חזותיים
Scaling GUI Agents with Visual State Transitions
State Transition Pretraining (STP) הוא שיטה חדשה להגברת GUI agents. השיטה מאפשרת למודל ללמוד ממעברי מצב חזותיים ולשפר את ביצועיו. המחקר הראה ש-STP משפר את הביצועים של המודלים ב-GUI scenarios.
תקציר מקורי באנגליתarXiv:2607.24112v1 Announce Type: new Abstract: We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predicting actions from state changes) and forward dynamics (predicting next states from current states and actions). This optimization equips the model with better action-grounded visual representations and an internal world model of GUI dynamics. When subsequently fine-tuned on trajectories with task instructions, our STP-trained models consistently outperform baselines trained solely via direct trajectory fine-tuning across agent benchmarks in both desktop and mobile GUI scenarios (AgentNetBench, AndroidControl, and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית