כתבה
arXiv cs.AI ·
CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling
תקציר מקורי באנגליתarXiv:2609.08686v1 Announce Type: cross Abstract: Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge access. Recent long-context language models can perform chaptering from textualized video inputs, but they remain costly and brittle for content-dense lecture videos with long transcripts, smooth topic transitions, and detailed chapter outputs. A scalable segment-then-caption paradigm reduces this cost, but introduces two new challenges: boundary error propagation and fragmented cross-chapter context. We propose \textbf{CausalChapter}, an intervention-inspired framework for long-video chaptering that estimates prediction-level influence through lightweight masking and removal interventions. For boundary localization, our Local Dependen
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית