יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

שיתוף שיעורים SFT ברחבי היישור, אורגניזמים מודליים ומודלים צעצוע

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
חוקרים בודקים האם שיעורים מאורגניזמים מודליים ומודלים צעצוע יכולים לשמש בהיישור. הם מוצאים שאימון על סיבות להתנהגות יכול לשפר את ההתנהגות. הם גם מוצאים שאימון על פלטים שנכתבו על ידי מודל אחר יכול לפגוע ביכולות.
תקציר מקורי באנגליתarXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. When projects share a goal, we should test whether lessons learned from one area transfer to the other areas. We study three such transfers, each taking a lesson developed in one SFT setting and testing it in another. First, we port a lesson about behavior generalization from alignment training into toy models. Training on the reason for a behavior, as in Teaching Claude Why, can make the behavior generalize better than training on examples of the behavior alone. Second, we port a lesson about capability preservation from model organisms i
קרא במקור המקורי