יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

The Geometry of Harmfulness in Multi-Turn Attacks

תקציר מקורי באנגליתarXiv:2609.38389v1 Announce Type: cross Abstract: Large language models (LLMs) remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn attacks, and why single-turn defenses are less effective in multi-turn settings. This work investigates how the geometry and temporal dynamics of harmfulness and refusal representations evolve across multi-turn attacks. We analyzed hidden-state representations from three instruction-tuned LLMs (Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Gemma-2-9B-it) using three multi-turn attack frameworks (Crescendo, ActorAttack, and X-Teaming), and examined representation behavior across conversation turns, model layers, a
קרא במקור המקורי