יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

How Inefficient Is Natural Gradient Descent? From Exact Optimality to \Theta ( \sqrt{ \log d } ) Divergence

תקציר מקורי באנגליתarXiv:2610.07228v1 Announce Type: cross Abstract: Natural gradient descent (NGD) underlies common methods in ML. For dually flat families, idealized NGD on the forward Kullback--Leibler objective follows the mixture geodesic which is often longer than the shortest Fisher--Rao path. We quantify this overhead by the inefficiency ratio \(R \ge 1\), the Fisher length of the mixture geodesic divided by the Fisher--Rao distance, and bound its supremum over endpoint pairs as a function of the parameter dimension \(d\). A tensor criterion identifies the regime (I) families, with \(R=1\) everywhere: exactly those with quadratic potential or dimension one, such as fixed-covariance Gaussians. For non-quadratic families, we prove two further regimes: (II) bounded third-order skewness plus finite Fishe
קרא במקור המקורי