כתבה
arXiv cs.AI ·
הפעלות טריינינג, ex post facto: חיבור אמונות המודל לאורך נקודות צומת
Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints
אנשי מחקר חוקרים טכניקות להפעלות טריינינג סינתטיות, כדי לשפר את האמונות של המודל. הם חוקרים את השפעתן על המודל, ומציעים פתרון חדש.
תקציר מקורי באנגליתarXiv:2610.00767v1 Announce Type: cross Abstract: Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run before its effect can be measured, making iteration slow and expensive. Common practice instead applies SDF to an already post-trained model. This is known to leave artifacts and degrade capabilities, and, as we show, it makes the model treat fabricated entities unrelated to the documents as real
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית