יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

תלוי במאגר הנתונים: כאשר ניבויי מודל קידוד מוחי עולים על גבי השלד הוויזואלי שלהם לזכרון וידאו

It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability
מודלים של קידוד מוחי מנבאים את זכרון הווידאו. מחקר זה בודק אם ניבויי המודלים האלה יכולים לשמש כתכונה מועילה לצורך חיזוי זכרון וידאו. התוצאות מראות כי הניבויים עולים על השלד הוויזואלי במקרים מסוימים, אך לא בכולם.
תקציר מקורי באנגליתarXiv:2607.16292v2 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge. We ask whether their predicted responses, obtained with no scanner, are a useful feature lens for a downstream human-behavior task: forecasting the memorability of short videos. We project each clip into TRIBE v2's predicted cortical space and forecast short-term memorability with ridge regression, against a matched control: the model's own V-JEPA2 visual backbone taken before the brain projection. The answer is dataset-dependent, and cleanly so. Within Memento10k the backbone wins (Spearman 0.594 vs 0.544 for the brain projection); within VideoMem the brain projection wins (0.415 vs 0.368, delta +0.047,
קרא במקור המקורי