יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Fed-GRPO: אופטימיזציה פדרטיבית

Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization
Fed-GRPO הוא שיטה חדשה לאופטימיזציה פדרטיבית של מודלים. היא מאפשרת אימון משותף של מודלים ללא חשיפת נתונים רגישים. השיטה משתמשת באותות רווח כדי לנווט את תהליך האימון.
תקציר מקורי באנגליתarXiv:2610.11502v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL), particularly through Group Relative Policy Optimization (GRPO). However, existing GRPO methods assume centralized access to training data, which may not hold in practice due to privacy or regulatory constraints. To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication. Fed-GRPO contains three reward-signal-driven mechanisms: (i) \emph{signal-weighted
קרא במקור המקורי