יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SAFESHIELD: תשתית לארגון החלטות לביטחון פריסה-זמן של מודלי שפה קטנים

SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models
תשתית חדשה לביטחון פריסה-זמן של מודלי שפה קטנים. SAFESHIELD מאפשרת ארגון החלטות לביטחון פריסה-זמן של מודלי שפה קטנים.
תקציר מקורי באנגליתarXiv:2610.07276v1 Announce Type: cross Abstract: Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capable mechanisms for these functions, but offer limited guidance on how the safety decisions they produce should be explicitly organized, coordinated, and audited. We formulate deployment-time safety as a decision-organization problem with two elements: responsibility-oriented decomposition of safety decisions and explicit coordination among them. We instantiate this formulation in SAFESHIELD, a deployment-time safety system for small language models that organizes four recurring decision responsibilities (admission,
קרא במקור המקורי