יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

קריאה, לא התערבות: ניצול לוגיטים של ראוטרים לבטיחות רב-מודאלית

Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models
חוקרים גילו כי לוגיטים של ראוטרים יכולים לשמש כאותות אבחוניים לבטיחות רב-מודאלית. הם פיתחו כלי לזיהוי בקשות לא בטוחות במודלים Qwen3-VL ו-Kimi-VL.
תקציר מקורי באנגליתarXiv:2610.07774v1 Announce Type: new Abstract: Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provid
קרא במקור המקורי