יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

WebFovea: כשהמודל נכון אבל הלחיצה טועה — טיוב יציב לאג'נטי ראייה-בסיסיים באתרי אינטרנט חיה

WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites
WebFovea היא אג'נט ראייה-בסיסי שהגיעה למקום השני בתחרות WebRetriever 2026. האג'נט פועלת על אתרי אינטרנט חיה ומספקת תשובות זמינות. האג'נט משתמשת במודלי גירוי גדולי-הרבה-לשוניים (LLM) כדי לבצע פעולות ראייה-בסיסיות כגון ניתוח תמונות, הגדרת רצפים ועוד. האג'נט פועלת באמצעות חיבור לאתרי אינטרנט ומספקת תשובות זמינות. האג'נט נחשבת לאחת האג'נטים המתקדמות בתחום הראייה-בסיסיים.
תקציר מקורי באנגליתarXiv:2610.03036v1 Announce Type: cross Abstract: We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but not sufficient. The model's decisions reach the browser through the harness, the code between the model and the page. At every step, four things must go right: the model's reply must be parsed into the intended action, the action must take effect on the page, the result must be reported back accurately,
קרא במקור המקורי