כתבה
arXiv cs.AI ·
ערכים כסגנון: פיצול ערכים מסמנטיקה עם תערובת אחד-כיוון לשליטה נמוכה ב-LLM
Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering
במאמר זה, המחברים מציגים פתרון לשליטה ב-LLM שמשמר את התחום והעובדות, ומשנה רק את הערכים. הם מציגים תערובת אחד-כיוון שמאפשרת זאת, ומדגימים את יעילותה במבחן על LLaMA-3.1-8B.
תקציר מקורי באנגליתarXiv:2609.39701v1 Announce Type: new Abstract: Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-g
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית