יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בדיקת שיטות ניקיון-קל-והשאיבה של קטעי-קליפ עבור תחושת-אורך-מערכת-מדיה

Evaluation of MLLM-Agnostic Plug-and-Play Keyframe Selection Methods for Long Video Understanding
בדיקה של שיטות ניקיון-קל-והשאיבה של קטעי-קליפ עבור תחושת-אורך-מערכת-מדיה. המאמר עוסק בשיפור יכולת ההבנה של MLLMs לסרטוני-וידאו-ארוכים.
תקציר מקורי באנגליתarXiv:2609.13250v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) cannot process every frame of a long video because of limitations in visual-token and computational budgets. Three primary approaches have been proposed to enhance their long-video understanding capabilities: (i) Retraining an MLLM on a large video corpus and/or extending its input length; (ii) Training an adapter for a specific MLLM that takes the entire video and the query as input and selects the most relevant video frames; and (iii) Developing a training-free, plug-and-play (PaP) adapter that is MLLM-agnostic. We refer to the third approach as PaP keyframe selection. A PaP method may use only candidate video frames without considering the query, or it may use both candidate video frames and the q
קרא במקור המקורי