כתבה
arXiv cs.AI ·
VAMR: תיאום רציונלי מרוב-שאלות להבנה יעילה של וידאו ארוך
VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding
VAMR הוא תיאום רציונלי של שאלות רבות להבנה יעילה של וידאו ארוך. הוא משתמש ב-Gemini ומציע תיאום רציונלי של שאלות רבות להבנה יעילה של וידאו ארוך.
תקציר מקורי באנגליתarXiv:2610.11171v1 Announce Type: cross Abstract: Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence jointly and progressively build a shared understanding that supports the complete question set. We introduce \textbf{VAMR} (\textbf{V}ideo \textbf{A}gent for \textbf{M}ulti-Question \textbf{R}easoning), which coordinates all questions about a video through one shared tool-use trajectory. At each round, a persistent policy model can invoke tools for one or more unresolved questions and submit answers for questions with sufficie
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית