כתבה
arXiv cs.AI ·
VideoScout: למידת חקירה פעילה עם קצב תכנון אדפטיבי
VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
VideoScout הוא סוכן רב-תורי שלומד חקירה פעילה עם קצב תכנון אדפטיבי להבנת וידאו ארוכים. המודל מאפשר גישה לתוכן וידאו רב יותר תוך שמירה על עומק ניתוח התוכן. הוא מתאמן על מערך נתונים גדול ומראה ביצועים חזקים בהשוואה למודלים אחרים.
תקציר מקורי באנגליתarXiv:2609.15606v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on short video understanding yet remain limited on long videos due to the limited visual context window. Prevailing approaches rely on uniform frame sampling or recent coarse-to-fine agentic zooming, both of which struggle to localize sparse, decisive evidence in sufficiently long videos. We formulate long video understanding as a \textbf{Sequential Evidence Acquisition (SEA)} problem, in which an agent reads the video turn by turn along the temporal axis, deciding at each turn how fast to watch, what evidence to retain, when to revisit uncertain segments, and when to stop and answer. Inspired by this view, we propose \textbf{VideoScout}, a multi-turn reasoning agent
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית