כתבה
arXiv cs.AI ·
לפני שהיא נעלמת: חיזוק התייחסות זמנית בזמן הערכת המודל
Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
אנו חוקרים את החולשה התקיפה של המודלים הגדולים לשפה וידאו (VideoLLMs) בתחום ההסתברות הזמנית. המאמר עוסק בפיתוח שיטה לחיזוק התייחסות זמנית בזמן הערכת המודל. ניתן למצוא את הקוד ב-GitHub.
תקציר מקורי באנגליתarXiv:2610.01595v1 Announce Type: cross Abstract: Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $\tau_l$, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית