יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

לא רואה את הסכנה: ציון נייטרוסטיות של תצפיות ראייה-לשון קפואות

It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank
מודלי ראייה-לשון קפואות עשויים לא לזהות סכנות כפי שמצופה.
תקציר מקורי באנגליתarXiv:2610.09517v1 Announce Type: cross Abstract: Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. VLM-based methods have reported gains in driving and safe-RL benchmarks by converting image-text similarity into rewards, costs, or confidence weights. Such signals promise to reduce reliance on manually designed feedback. They may also reflect prompt structure, embedding geometry, or camera viewpoint, leaving their safety meaning unverified. To address this gap, we present a controlled evaluation of a frozen CLIP prompt-margin
קרא במקור המקורי