יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

תקציר מקורי באנגליתarXiv:2606.09701v2 Announce Type: replace-cross Abstract: Language model safety must continually adapt to evolving attacks. Recent works have demonstrated that reinforcement learning can be used to train stronger attacker and defender models in tandem by applying PPO-style self-play and DPO-style online preference optimization. In this work, we explore the efficacy of GRPO in this setting. Co-training can be challenging because it requires jointly optimizing multiple properties of both the attacker and defender. We therefore shape model outputs using multiple LLM judge-based reward channels and compute advantages with GDPO, which prevents any single channel from dominating. Our method uses a curriculum that progresses from attacker-only single-turn and multi-turn training to co-training, w
קרא במקור המקורי