יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

חיזוי ותיקון קריסת מיזוג במודלי שפה גדולים

Predicting and Repairing Merge Collapse in Large Language Models
חוקרים פיתחו שיטה לחיזוי ותיקון קריסת מיזוג במודלי שפה גדולים. השיטה, PRISM, משפרת את היכולת למזג מודלים מבלי לאבד ביצועים. המחקר נערך על 22 קונפיגורציות מיזוג מארבע משפחות מודלים.
תקציר מקורי באנגליתarXiv:2610.03199v1 Announce Type: new Abstract: Large language models fine-tuned from a shared base can be merged by averaging their task vectors, but some merges collapse far below the base model, and common merge operators give no warning before evaluation. We show that one statistic of the specialists' task vectors both predicts this collapse and calibrates its repair. The power that averaging removes equals the variance of the task vectors across specialists, our measure of interference. Under a working noise model, the disturbance that a merge injects grows with the merge coefficient and with interference, yielding a pre-merge score. In our experiments on twenty-two merge configurations from four model families, only destructive merges exceed a threshold on this score. We find that st
קרא במקור המקורי