כתבה
arXiv cs.CL ·
ניבוי תכונות כלליות של תכנון התאמה
Predicting Alignment Generalization with Value Representations
במאמר זה, חוקרים פותחים טכניקות לניבוי תכונות כלליות של תכנון התאמה ב-LLMs. הם מציעים דרך לפתח תכונות כלליות של תכנון התאמה, ומדגימים את יעילותה במספר תחומים.
תקציר מקורי באנגליתarXiv:2610.12410v1 Announce Type: new Abstract: LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית