יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

שמירת בטיחות מודלי שפה גדולים: COLAGUARD

Robust and Efficient Guardrails with Latent Reasoning
COLAGUARD משפר את שמירת הבטיחות של מודלי שפה גדולים באמצעות סיבוכיות לא-מודעת. המודל COLAGUARD עובד על 10 סטים שונים של פרסומים ותגובות, ומשפר את ה-F1 המקדם ב-8.24 נקודות יותר מ-Llama Guard 3. COLAGUARD גם מציע עדיפות של 12.9X בקצב העברה ו-22.4X בחיסכון בתגובות.
תקציר מקורי באנגליתarXiv:2605.29068v2 Announce Type: replace-cross Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrails typically rely on single-pass classification or, more recently, distilled reasoning. Reasoning-based guardrails significantly outperform classification-only baselines, but they incur substantial query latency and token overhead that make them impractical for highthroughput deployment. To address this challenge, we propose COLAGUARD, a guardrail model that transfers multi-step safety reasoning into a continuous latent space through a stage-wise training curriculum, enabling direct hidden-state propagation at inference. Evaluated on ten prompt- and response-moderation settings spann
קרא במקור המקורי