כתבה
arXiv cs.CL ·
ForeSight: שיפור מעקב סיכונים
ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation
ForeSight הוא כלי לשיפור מעקב סיכונים על ידי ניתוח אותות בטיחות מוקדמים. הוא משתמש במודלים גדולים של שפה כדי לזהות תוכן מסוכן. הכלי מוכיח יעילות בחמישה מבחני בטיחות ושני מודלים יעד.
תקציר מקורי באנגליתarXiv:2609.13737v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards operate at the input, output, or streaming-generation stages, while early-risk methods that rely on surface tokens or output logits may suffer from weak initial signals, and internals-based detectors using dense representations may retain highly entangled and redundant safety-irrelevant information. It therefore remains unclear whether the earliest post-generation hidden states already contain reliable signals about final-response harmfulness. To address this gap, we propose ForeSight, a first-token output-risk forecasting framework that distills weak and redundant early safety signals into com
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית