יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

TRACE: למידת פאטץ' בטיחות מבוססת מסלול

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
TRACE הוא כלי ללמידת פאטץ' בטיחות עבור מודלי שפה גדולים. הוא מאפשר לשחזר את הבטיחות של המודל ללא צורך באימון מחדש. TRACE נבדק על שישה בנצ'מרקים ושני מודלים, והשיג תוצאות מרשימות.
תקציר מקורי באנגליתarXiv:2607.16242v1 Announce Type: new Abstract: Fine-Tuning-as-a-Service (FTaaS) platforms let users train large language models (LLMs) on customized tasks, but this pipeline could erode models' safety alignment. In practice, service providers need to recover models' safety without re-running full alignment, or destroying the utility gained from customized tasks. A line of existing work refers to model parameter merging, which adds a safety patch on the fine-tuned model parameters to shift the model away from unsafe tendencies. However, this merging-based paradigm is fundamentally bottlenecked by task-safety update entanglement: downstream task updates and the safety patch often overlap in their dominant directions, so the merge strength is intrinsically hard to calibrate. If the safety ve
קרא במקור המקורי