כתבה
arXiv cs.LG ·
Omni-Streaming Thinking
תקציר מקורי באנגליתarXiv:2609.15128v1 Announce Type: new Abstract: Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasoning can keep relaying it even after audio contradicts it. We call this failure premature cross-modal commitment. We propose Omni-Streaming Thinking (OST), which generates structured outputs that include evidence observed so far, forecasts of future evidence, and claims based on this evidence. Each claim is initially marked as pending and linked to a future verification interval. Audio and visual evidence are stored separately, and OST checks a claim against the evidence
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית