יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

קצרות בזנב: הפחתת תופעות דווקא על ידי צמצום ספקטרלי של עדכוני הטמעה

Shortcuts in the Tail: Debiasing via Post-Hoc Spectral Compression of Fine-Tuning Updates
במאמר זה, המחברים מציעים פתרון לבעיית הטמעה של LLMs. הם מציעים לצמצם את הטאי של ה-LM על ידי צמצום ספקטרלי של העדכונים של ה-LM. הם מציגים תוצאות ניסויים שמדגימות את יעילות הפתרון.
תקציר מקורי באנגליתarXiv:2606.07596v2 Announce Type: replace Abstract: Fine-tuning often introduces spurious correlations alongside task knowledge, causing systematic failures on underrepresented groups. Existing mitigations require retraining, group labels, or curated counterfactual data. We show a simple post-hoc intervention reduces shortcut reliance without any of these: truncating the tail of the SVD of $\Delta W = W_\mathrm{ft} - W_\mathrm{base}$ reduces the spurious-group gap while preserving task accuracy. Across three instruction-tuned models ($0.5$B--$7$B) and four classification benchmarks, top-$k$ truncation reduces the gap on every cell at $<2$ pp accuracy loss, by up to $5\times$ on CivilComments. We propose this works because the shortcut response sits in the tail of the singular ordering of $
קרא במקור המקורי