יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בקבוטלת הוויזואלית: התאמה דלילה של MLLM לקרקע וידאו מרחבית-זמנית

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding
חברות וידאו גדולות עובדות עם מיליוני העלאות בשעה, וזקוקות למערכות איכות שיכולות לאתר היכן ומתי הפרות מדיניות קורות. המחקר מציג שיטות אימון לסגירת הפער בין תנאי האימון לתנאי הפריסה, ומוצא כי הפער הוויזואלי הוא המגבלה העיקרית. התאמה של שלוש שכבות ViT אחרונות, 4% מהפרמטרים, משיגה 68.8% דיוק זמני, ועוקפת מודל 8B שאינו מאומן עם קלט צפוף.
תקציר מקורי באנגליתarXiv:2607.24570v1 Announce Type: cross Abstract: Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing every frame is infeasible at scale, so systems are constrained to sparse inputs of 8 to 16 frames per video. Yet state-of-the-art multimodal large language models (MLLMs) are pretrained on dense sequences of hundreds of frames, creating a fundamental mismatch between training and deployment conditions. This mismatch causes severe performance collapse: the Qwen3-VL 8B model drops from 56.0% to 22.3% temporal mIoU when frames are reduced to 16, a 60.2% relative degradation. We present a systematic empirical study of training strategies to close this gap for spatial-
קרא במקור המקורי