כתבה
arXiv cs.LG ·
RAPTOR: Role-Aware Private Training for Mixture-of-Experts
RAPTOR - תורת רפורמינג פרטית למודלי MoE, המטפלת בבעיות פרטיות באמצעות שיטת RAPTOR
תקציר מקורי באנגליתarXiv:2609.05770v1 Announce Type: new Abstract: Differentially private (DP) fine-tuning methods treat sparse Mixture-of-Experts (MoE) models as a single dense block, ignoring that shared layers see all data while experts only see routed records. We identify and formally characterize three resulting failure modes: global clipping suppresses expert gradients, batch-level normalization dilutes sparse expert updates, and fixed privacy noise degrades signal-to-noise ratio on low-load experts. We introduce RAPTOR - a Role-Aware Private Training framework, which alternates shared and expert optimization and targets each failure directly, using expert-specific clipping and noise together with a public expected-owner denominator and a count-independent update schedule that avoids conditioning on pr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית