כתבה
arXiv cs.CL ·
שמירת בטיחות LLMs: COLAGUARD - גישה רציפה למחשבה
Robust and Efficient Guardrails with Latent Reasoning
COLAGUARD - גישה רציפה למחשבה: מודל שמטפל בבטיחות LLMs באמצעות חלל רציף. המאמר מציג פתרון חדשני לבטיחות LLMs, COLAGUARD, שמטפל בבטיחות באמצעות חלל רציף. זה מאפשר תפעול יעיל ומהיר יותר של LLMs בעת הפעלה.
תקציר מקורי באנגליתarXiv:2605.29068v2 Announce Type: replace-cross Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrails typically rely on single-pass classification or, more recently, distilled reasoning. Reasoning-based guardrails significantly outperform classification-only baselines, but they incur substantial query latency and token overhead that make them impractical for highthroughput deployment. To address this challenge, we propose COLAGUARD, a guardrail model that transfers multi-step safety reasoning into a continuous latent space through a stage-wise training curriculum, enabling direct hidden-state propagation at inference. Evaluated on ten prompt- and response-moderation settings spann
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית