כתבה
arXiv cs.LG ·
התערבויות באימון קדם-למידה: גידול תפיסות המודל בין נקודות דיווח
Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints
התערבויות באימון קדם-למידה שמקטינות דחייה ממשית. פיתוח טכנולוגיות להפחתת דחייה ממשית באימון קדם-למידה. טכנולוגיות להפחתת דחייה ממשית באימון קדם-למידה.
תקציר מקורי באנגליתarXiv:2610.00767v1 Announce Type: new Abstract: Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run before its effect can be measured, making iteration slow and expensive. Common practice instead applies SDF to an already post-trained model. This is known to leave artifacts and degrade capabilities, and, as we show, it makes the model treat fabricated entities unrelated to the documents as real,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית