כתבה
arXiv cs.LG ·
מה נפקד באימון פוסט-טריינינג? קריסה תקנית ואובדן של יכולת להנחות בהקשר
What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives
אימון פוסט-טריינינג גורם לירידה ביכולת המודל להתאים למידע בהקשר. ניסויים מומחשים זאת.
תקציר מקורי באנגליתarXiv:2610.02614v1 Announce Type: new Abstract: AI models serving a heterogeneous population must act on the principles appropriate to each user and context. While post-training has been shown to narrow the views large language models express, prior work has focused on default behavior rather than the ability to adapt to in-context information. We show that post-training also degrades a model's ability to be steered in-context toward perspectives it was not trained to favor. In controlled experiments, we fine-tune models toward one side of cultural-value disagreements and evaluate checkpoints throughout training. The trained side becomes increasingly dominant in ordinary use, while the ability to recognize and faithfully enact the opposing view declines. These findings point to a tension b
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית