כתבה
arXiv cs.CL ·
CineSubBench: בדיקת LLMs בסרטים
CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles
CineSubBench הוא בנק אבטיפוס לבדיקת מודלי שפה גדולים בהבנת סרטים. הוא מכיל 1,012 סרטים עם כיתובים בשש שפות. המחקר בודק יכולות של מודלים כמו LLaMA.
תקציר מקורי באנגליתarXiv:2609.36218v1 Announce Type: new Abstract: Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce CineSubBench, a benchmark for evaluating long-context film understanding from multilingual movie subtitles. A subtitle track represents a film as thousands of short, temporally ordered utterances from which models must reconstruct characters, relationships, events, causal progression, and themes without explicit scene or event structure. CineSubBench contains 1,012 films with complete subtitle coverage in six languages, yielding 6
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית