יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

חיזוי תכונות כלליות של תכנון התאמה

Predicting Alignment Generalization with Value Representations
במאמר זה, חוקרים חוקרים את חיזוי תכונות כלליות של תכנון התאמה, כלומר, חיזוי איך טיפול נוסף של מודל להתאים לערך מסוים ישפיע על התנהגותו במגוון רחב של ערכים שאינם ידועים.
תקציר מקורי באנגליתarXiv:2610.12410v1 Announce Type: cross Abstract: LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques
קרא במקור המקורי