כתבה
arXiv cs.LG ·
Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?
תקציר מקורי באנגליתarXiv:2610.07444v1 Announce Type: new Abstract: A GUI agent decides which action to take and where to take it; we ask how a small grounding model should receive the action type. Fine-tuning Qwen2-VL-2B with LoRA on Android in the Wild, we compare a flat baseline with five ways of supplying the type under matched data, compute, and decoding: an auxiliary loss, a hard-routed action word, an additive learned embedding, a prepended learned token, and the type written into the prompt. With five seeds, an episode-clustered bootstrap, and seed-level paired tests, the ranking on a mixed stream is clear: the auxiliary loss, the additive embedding, and the prompt word each gain five to seven hit@0.10 points over the baseline, while hard routing and the prepended token are not distinguishable from it
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית