כתבה
arXiv cs.AI ·
מבחן ושיפור: תקן ושיפור תקיפה-עם-וידאו למודלי יצירת וידאו
From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
במאמר זה, נוצרה במבחן ושיפור תקיפה-עם-וידאו למודלי יצירת וידאו. הבמבחן, שנקרא VWG-Bench, כולל 9 תפיסות ו-38 תפיסות דקות. המאמר גם מציג פקטור של Vid-PRE, שהוא פקטור של Vid-PRE.
תקציר מקורי באנגליתarXiv:2609.11242v1 Announce Type: cross Abstract: Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often conflating visual quality with cognitive correctness. We introduce VWG-Bench (Video World Generalist Benchmark), a comprehensive benchmark spanning 9 reasoning dimensions and 38 fine-grained tasks. To enable precise diagnosis, we design a three-level VLM-as-Judge protocol that independently assesses video-level fluency, task-level rule adherence, and sample-level goal realization. Evaluations of leading models reveal a striking g
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית