כתבה
arXiv cs.AI ·
StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions
תקציר מקורי באנגליתarXiv:2609.01081v2 Announce Type: replace-cross Abstract: Large language models often answer the same multiple-choice question inconsistently when it is posed under support-oriented and elimination-oriented framings. We investigate whether these discrepancies arise from different internal representations induced by the two framings. We introduce a dual-framing protocol with minimally varied prompts that use either support- or elimination-oriented framing while keeping the evaluation target fixed. To probe the internal computation, we append an untrained special token, [STATE], and treat its residual-stream activation as an intervention interface. Across both models, the two framings induce separable [STATE] activations concentrated in intermediate layers. Swapping these activations between
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית