כתבה
arXiv cs.AI ·
תחזית התאמה: חיזוי חוסר התאמה מדגם הכשרה
Alignment Forecasting: Predicting Misalignment From Training Data
חיזוי חוסר התאמה מדגם הכשרה. חוקרים הציגו שיטה לחזות כיצד דגם לשוני עלול להיות חוסר-תאום. השיטה, שנקראת Alignment Forecasting, עושה שימוש באלגוריתם של LLM כדי לבדוק את הסיכויים של דגם להיות חוסר-תאום. החוקרים טענו כי Alignment Forecasting יכול לסייע למפתחי LLM לזהות ולתקן חוסר-תאום לפני שהדגם נכנס לשימוש.
תקציר מקורי באנגליתarXiv:2609.35805v2 Announce Type: replace-cross Abstract: Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 3
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית