יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מהו התפנית? בדיקת עקביות זמנית במודלים ויז'ואליים-לשוניים

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models
חוקרים פיתחו TimeCatch, כלי לבדיקת עקביות זמנית במודלים ויז'ואליים-לשוניים. המחקר מראה פער ניכר בין גילוי חריגות ברמת הפריימים לבין גילוי חריגות זמניות. בניגוד לבני אדם, המודלים הוויז'ואליים-לשוניים מתקשים לזהות חריגות זמניות.
תקציר מקורי באנגליתarXiv:2608.23474v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection problem, providing a simple and controlled evaluation that directly tests sensitivity to temporal consistency. We introduce TimeCatch, where temporal anomalies are created by swapping consecutive frames and frame-level anomalies by replacing a frame with Gaussian noise. Models are evaluated on anomaly detection and localization tasks across four synthetic and real-world datasets, alongside a human study. Our evaluation reveals a substantial gap between frame-level and temporal anomaly detection.
קרא במקור המקורי