כתבה
arXiv cs.LG ·
למידה פדרטיבית סקלרית למורכבת רגולטור לינארי-קוואדרטי
Scalar Federated Learning for Linear Quadratic Regulator
אלגוריתם פדרטיבי יעיל ללמידה של שליטה במורכבת רגולטור לינארי-קוואדרטי. המאמר מציג פתרון ללמידה פדרטיבית של שליטה במורכבת LQR, עם יעילות תקשורת גבוהה. האלגוריתם, הנקרא ScalarFedLQR, מבנה על מנגנון גרדיאנט משופר, שבו כל סוכן מקשר רק עם סקלר של הגרדיאנט המקומי. השרת מאגדת את המסגרות הסקלריות כדי לשחזר כיוון ירידה גלובלי, וכך קוטעת את התקשורת העל-סוכן מ-O(d) ל-O(1), עם ניחוש טוב.
תקציר מקורי באנגליתarXiv:2604.05088v2 Announce Type: replace-cross Abstract: We propose ScalarFedLQR, a communication-efficient federated algorithm for model-free learning of a common policy in linear quadratic regulator (LQR) control of cooperative agents. The method builds on a decomposed projected gradient mechanism, in which each agent communicates only a scalar projection of a local zeroth-order gradient estimate. The server aggregates these scalar messages to reconstruct a global descent direction, reducing per-agent uplink communication from O(d) to O(1), independent of the policy dimension. Crucially, the projection-induced approximation error diminishes as the number of participating agents increases, yielding a favorable scaling law: larger fleets enable more accurate gradient recovery, admit large
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית