יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ספק תוצאות רעות עוזר לזיהוי תיקוני תצוגה

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates
נמצא כי תיקוני תצוגה של מודלי שפה גדולים יכולים להכיל ראיה להתנהגות רעה. זה עשוי לאפשר אפקטיביות גבוהה יותר באודיטינג ובפיקוח.
תקציר מקורי באנגליתarXiv:2610.07518v1 Announce Type: new Abstract: Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates? We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space. Using a reference geometry defined by pure harmful-compliance, safety-targeted, and benign-utility SFT, we find that a checkpoint-level coordinate s_H tracks controlled harmful-objective composition with Spearman correlations of 0.986-0.992 across four 7-8B backbones, with the same ordering persisting
קרא במקור המקורי