כתבה
arXiv cs.LG ·
משפט הנגזרת הפיזית לאימון סטוכסטי
Physical policy gradient theorem for in situ stochastic-adjoint training
חוקרים מציגים אלגוריתם חדש לאימון מודלים באמצעות נגזרות סטוכסטיות. האלגוריתם מאפשר אימון ישיר מתוך מדידות, ללא צורך בניסויים נפרדים. המחקר מראה יישום מוצלח ברשתות רונולוציות לא ליניאריות.
תקציר מקורי באנגליתarXiv:2609.05808v1 Announce Type: cross Abstract: In situ adjoint training extracts parameter gradients directly from measurement, but has so far been limited to reciprocal or restricted systems. Here, we introduce the physical counterpart of the policy gradient theorem: a stochastic-adjoint gradient estimator that lifts these constraints by trading reciprocity for nondegenerate diffusion. As validation, we train a nonlinear resonator network, whose own dynamics supply the policy, against antagonistic temporal modulations with gradients from measured stochastic trajectories alone, without finite differences or a separate adjoint experiment.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית