יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SCOPE-OPSD: תחום-OPSD: תת-מרחבי הפריבילג'ד של SCOPE-OPSD להתנהגות-OPSD

SCOPE-OPSD: Fisher-Conditioned Privileged Subspaces for On-Policy Self-Distillation
SCOPE-OPSD מציע תת-מרחבי הפריבילג'ד של SCOPE-OPSD להתנהגות-OPSD. המחקר חוקר את האפשרות לשימוש בתחום הפריבילג'ד כדי לשפר את ההתנהגות של OPSD. התוצאות המוצגות במחקר תומכות בשימוש בתת-מרחבי הפריבילג'ד של SCOPE-OPSD.
תקציר מקורי באנגליתarXiv:2609.12579v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) scores student-generated prefixes with a solution-conditioned self-teacher, yet transfers supervision only through next-token probabilities. We ask whether the aligned final-layer discrepancy offers a useful second channel, and how to test that channel without confusing its geometry with auxiliary strength. SCOPE-OPSD projects the privileged teacher-student residual onto a frozen rank-64 factor estimated from residual covariance and language-model-head Fisher sensitivity. It reuses the forwards already required by OPSD and adds neither rollouts nor inference-time modules. A matched Random control preserves the structured factor's rank and nonzero spectrum and uses per-arm gradient-RMS calibration, isolating
קרא במקור המקורי