כתבה
arXiv cs.LG ·
Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering
תקציר מקורי באנגליתarXiv:2609.39701v1 Announce Type: cross Abstract: Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית