כתבה
arXiv cs.LG ·
מעבר למצב-במקום ואפס-מודל: ניתוח ושיפור
Decodable In-Context State and Model Output Across Training
במאמר זה, נחקרה האפשרות לקריאה של מצב-במקום ואפס-מודל במהלך תרגילי הכשרה של מודל Pythia. נמצא כי הדיוק של הקריאה עולה במהלך הכשרה, ושיטת הנחיה-מודדת יכולה לתקן טעויות.
תקציר מקורי באנגליתarXiv:2609.31401v1 Announce Type: new Abstract: Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish probe-correct errors with low and above-uniform model probability for the correct candidate. Oracle-target steering already repairs many early errors, but saved aggregates cannot separate target quality from intervention sensitivity. A held-out comparison of decoders trained on the final state or candidate log
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית