כתבה
arXiv cs.AI ·
CS-Guard: בדיקת גדרות למניעת יצירת קוד מזיק
CS-Guard: Benchmarking LLM Guardrails for Code Generation Security
CS-Guard הוא בנץ'מרק הראשון לבדיקת גדרות למניעת יצירת קוד מזיק. הוא בודק 9 גדרות ב-7 מודלים שונים, כולל GPT. התוצאות מראות כי הגדרות הנוכחיות מתקשות למנוע יצירת קוד מזיק.
תקציר מקורי באנגליתarXiv:2609.09798v1 Announce Type: cross Abstract: Large language models (LLMs) have been ex- ploited to generate malware, but the effective- ness of guardrails for code generation secu- rity remains unclear. We introduce CS-Guard, the first benchmark to systematically evalu- ate guardrails for code generation security. It covers 1) text-to-code generation with 1000 high-quality malware-generation prompts, 7 jailbreak attacks, and a novel fictional scenario attack (FSA) that embeds malicious intent in a legitimate fictional software-development sce- nario; and 2) code-to-code generation with 331 code prompts spanning code infilling, code completion, and code translation. We empiri- cally evaluate 9 guardrails across seven LLMs. We find that current guardrails perform poorly against maliciou
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית