יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

צפייה להתאמה: חיזוי פגמים בהתאמה מטרינינג נתונים

Alignment Forecasting: Predicting Misalignment From Training Data
חיזוי פגמים בהתאמה של מודלי לשון על ידי ניתוח נתוני הלמידה. המחקר מציג פלטפורמה לחיזוי פגמים בהתאמה של מודלי לשון, ומציע דרך לצפות ולמנוע פגמים בהתאמה של מודלי לשון.
תקציר מקורי באנגליתarXiv:2609.35805v1 Announce Type: new Abstract: Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets
קרא במקור המקורי