יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הפרדה בין התאמת ייצוג לביטוי אישיות במודלים שפה

K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models
חוקרים K/V-cache interventions במודלים שפה מסוג decoder-only, ומצאו הפרדה בין התאמת ייצוג לביטוי אישיות. המחקר בוחן 13 קונפיגורציות התערבות על מודל Llama-3.1-8B.
תקציר מקורי באנגליתarXiv:2609.11020v1 Announce Type: new Abstract: We study K/V-cache interventions -- transplanting a target-conditioned K/V trajectory into a source-persona generation -- as a structured surface for persona control in decoder-only language models. Across 13 intervention configurations applied to Llama-3.1-8B for a fixed source-to-target persona pair, we report two consistent dissociations between representation-level alignment and behavioral expression, plus a common failure under position perturbations. First, all layer-band K/V replacements (early, mid, late) achieve strong local V-space alignment (V-gap 0.91, 0.89, 0.84), but only mid-layer replacement (layers 9-20) combines substantial target-marker expression with preserved lexical diversity. Second, full and mid-layer replacement indu
קרא במקור המקורי