כתבה
arXiv cs.CL ·
מדוע הטכנולוגיה של וידאו עדיין יקרה? סקירה של מנגנוני יעילות-הפרעה ב-VideoLLMs
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
סקירה של מנגנוני יעילות-הפרעה ב-VideoLLMs, כולל תיאור של חידושים ב-VideoLLMs ובמודלי LLMs.
תקציר מקורי באנגליתarXiv:2609.10355v1 Announce Type: cross Abstract: Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a computation and memory cost that grows with frame count and context length, limiting deployment in real-time, mobile and resource-constrained settings. This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count. We analyze bottlenecks across frame sampling, modality encoding, connecto
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית