יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

VLX-VR: מודל סיבוב-וידאו מודע לאגנט

VLX-VR: An Agentic-Aware Video Reasoning Model
מודל סיבוב-וידאו חדש, VLX-VR, הציג תוצאות שיא בביצועיו. המודל, שנלמד באמצעות רכיבת למידה, יכול לקבל החלטות ולהחליט על פעולות בהתאם למידע שהוא קובע. המודל נלמד על ידי נתוני וידאו ומסלולי אגנטים, והוכיח יכולת גבוהה בביצועיו. המודל יכול לשמש בתחומים שונים, כגון זיהוי חפצים, זיהוי קול, ועוד.
תקציר מקורי באנגליתarXiv:2609.09985v1 Announce Type: new Abstract: Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and single-pass inference, limiting adaptive evidence acquisition when observations are incomplete, ambiguous, or conflicting. We present VLX-VR, an agentic-aware video reasoning model trained within a video reasoning framework defined by a Think--Memory--Observation loop. At each step, VLX-VR determines the needed evidence, invokes read_memory or write_memory, incorporates the returned Observation, and decides whether to continue or produce the task output. We train VLX-VR with multimodal data, including videos and agent trajectories, using reinforcement learning to learn e
קרא במקור המקורי