כתבה
arXiv cs.LG ·
בטיחות LLM: חקירת שפה 'שיכורה'
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
חוקרים בדקו את הבטיחות של מודלי LLM עם 'שפה שיכורה'. הם מצאו כי שימוש בשפה זו עלול לגרום לדליפות פרטיות ואיומים בטיחותיים. המחקר השתמש במודלים כמו LLaMA ו-LangChain.
תקציר מקורי באנגליתarXiv:2601.22169v2 Announce Type: replace-cross Abstract: Humans are susceptible to undesirable behaviours and privacy leaks under the influence of alcohol. This paper investigates drunk language, i.e., text written under the influence of alcohol, as a driver for safety failures in large language models (LLMs). We investigate three mechanisms for inducing drunk language in LLMs: persona-based prompting, causal fine-tuning, and reinforcement-based post-training. When evaluated on 5 LLMs, we observe a higher susceptibility to jailbreaking on JailbreakBench (even in the presence of defences) and privacy leaks on ConfAIde, where both benchmarks are in English, as compared to the base LLMs as well as previously reported approaches. Via a robust combination of manual evaluation and LLM-based eva
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית