יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בחירת נתונים מודולרית עם פישר-הדרכה להמשך-אימון של LLM גדול

Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models
בחירת נתונים מודולרית עם פישר-הדרכה להמשך-אימון של LLM גדול. ניתן לבחור נתונים שישפיעו על ה-LLM באופן יעיל ובעל תפוקה גבוהה.
תקציר מקורי באנגליתarXiv:2610.02593v1 Announce Type: new Abstract: Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chosen target-domain corpus can overwrite capabilities encoded in the pretrained checkpoint. Existing CPT practice either scores candidates with parameter-agnostic scalars such as perplexity, or mitigates forgetting by spending many extra general-domain replay tokens. Neither strategy directly asks how training on a candidate will move the model parameters. We show that loss-based selection causes the post-CPT Fisher diagonal to drift downward on exactly the high-Fisher coordinates the pretrained model had committed to
קרא במקור המקורי