יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בחינה של נטיות פרסונליות-תלויות במודלי שפה

Probing Persona-Dependent Preferences in Language Models
במאמר זה נחקרו נטיות פרסונליות-תלויות במודלי שפה. נמצא כי נטיות אלו נשמרות במודלי Gemma-3-27B ו-Qwen-3.5-122B.
תקציר מקורי באנגליתarXiv:2605.13339v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and prompting appear to influence much of their behaviour. But models can also adopt different personas which have radically different preferences. How is this implemented internally? Does each persona use its own preference representations, or are some representations shared? We train linear probes on residual-stream activations of Gemma-3-27B and Qwen-3.5-122B to predict revealed pairwise task choices, and identify a genuine preference vector: it tracks the model's preferences as they shift across a range of prompts and situations, and on Gemma-3-27B steering along it causa
קרא במקור המקורי