כתבה
arXiv cs.CL ·
LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception
תקציר מקורי באנגליתarXiv:2609.39938v1 Announce Type: new Abstract: Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framewo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית