יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

דינמיקת התפשטות של נזק: טרקטוריות של כוונת תקיפה במודלי שפה גדולים

Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models
נמצאה דינמיקת התפשטות של נזק במודלי שפה גדולים. טרקטוריות של כוונת תקיפה נמצאו לאורך עומק המודל. המחקר מציע פתרון לאבטחת מודלי שפה.
תקציר מקורי באנגליתarXiv:2609.13534v1 Announce Type: new Abstract: We identify \textbf{Harmfulness Propagation Dynamics (HPD)}: for harmful prompts, the projection of the last-token hidden state onto a learned harm direction rises monotonically with transformer depth, whereas benign prompts remain flat or oscillatory. This cross-layer signature reflects harmful intent as a \emph{progressively resolved} semantic property: surface form appears early, while pragmatic intent consolidates later, making the \emph{trajectory shape} more informative than any single-layer snapshot. Moreover, LDA-based harm directions, learned per layer, remain stable across random splits (pairwise cosine similarity $>0.97$), supporting the projection sequence as a reproducible structured signal. Building on HPD, we introduce \textbf{
קרא במקור המקורי