יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas

תקציר מקורי באנגליתarXiv:2605.09159v2 Announce Type: replace Abstract: Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has used such directions as static handles for behavioral steering. We instead treat them as dynamic signals: probes we can monitor and intervene on as reasoning unfolds. We use the term polylogue to denote the time series of alignments between persona vectors and hidden activations over the course of generation. Experiments across four open-weight models show that polylogue features contain predictive signal for correctness comparable to low-dimensional activation summaries, while remaining interpretable through their associated persona directions. They also suggest
קרא במקור המקורי