כתבה
arXiv cs.AI ·
בחירת הטובים ביותר בטקסט: תיאור פשוט של הכותרת המקורית
Selecting The Most Informative Tokens in Natural Language Autoencoders
בחירת הטובים ביותר בטקסט: תיאור פשוט של הכותרת המקורית. המחברים חקרו את השאלה האם ניתן לבחור את הטובים ביותר בטקסט כדי להבין את הסיכונים. הם חקרו 4.7 מיליון הסברים ומצאו שהסימנים מהסטרקטורה של השיחה נותנים הסברים רלוונטיים יותר.
תקציר מקורי באנגליתarXiv:2609.37040v1 Announce Type: cross Abstract: Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית