כתבה
arXiv cs.CL ·
האם ראיתי מספיק? דגימות-שפה-וידאו קודדות סימן לתוקפן של הראיות
Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness
דגימות-שפה-וידאו קפואות קודדות סימן לתוקפן של הראיות. נמצא כי דגימות-שפה-וידאו קפואות קודדות סימן לתוקפן של הראיות. נמצא כי דגימות-שפה-וידאו קפואות קודדות סימן לתוקפן של הראיות.
תקציר מקורי באנגליתarXiv:2610.08560v2 Announce Type: replace-cross Abstract: Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question re
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית