כתבה
arXiv cs.AI ·
Homomorphic Advantage Operator
Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints
Homomorphic Advantage Operator הוא כלי לייצוב למידת חיזוק תחת Fully Homomorphic Encryption. הוא מונע התפרקות של קירובים פולינומיים ב-FHE. ה-Operator משמש לתיקון ה-Bellman drift.
תקציר מקורי באנגליתarXiv:2610.02074v1 Announce Type: new Abstract: Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcement learning (RL) requires replacing non-linear operations with polynomial approximations, which diverge catastrophically due to a unique recursive error phenomenon known as the Bellman drift. This article introduces the Homomorphic Advantage Operator (HAO), a stabilization framework designed to prevent polynomial approximation divergence in FHE-based deep RL. HAO adapts the zero-mean centering projection from advantage-based valu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית