כתבה
arXiv cs.CL ·
בחירה, עיבוי, והשקעה מחדש: מחקר מומחש של אלוקציה של טוקן-מבט-עירובי ב-MLLMs לטווח ארוך
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
במחקר זה, נבדקה אלוקציה של טוקן-מבט-עירובי ב-MLLMs לטווח ארוך. התוצאות הראו שבחירה, עיבוי, והשקעה מחדש של הטוקן-מבט-עירובי יכולים לשפר את הדיוק של ה-MLLMs. המחקר גם חשף כי תכנות פגמים בבסיס הקוד של ה-MLLMs יכולים להשפיע על התוצאות.
תקציר מקורי באנגליתarXiv:2609.03820v1 Announce Type: cross Abstract: Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three long-video benchmarks, and two answering models. Selection is the largest single lever: on LongVideoBench's hour-long bin, eight query-selected frames
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית