כתבה
arXiv cs.AI ·
CHILLGuard: תשתית בטיחות למודלי שפה גדולים סיניים
CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment
מערכת בטיחות למודלי שפה גדולים סיניים, עם סקאלביליטי לבניית נתונים
תקציר מקורי באנגליתarXiv:2606.15396v2 Announce Type: replace-cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns. While existing LLM safety guardrails excel in English or multilingual settings, they lack adaptation to Chinese-specific regulatory policies, cultural context, and linguistic nuances, failing to support fine-grained risk classification for diverse deployment needs. In this paper, we introduce a 5-macro, 31-micro category fine-grained risk taxonomy for Chinese scenarios, and build CHILLGuard: a dedicated Chinese LLM content safety guardrail. To address the critical scarcity of high-quality annotated Chinese safety data, we propose a scalable multi-stage data construction pipeline: we expand multi-source corpus via retrie
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית