יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

גרידי לייר-ווייס טריינינג: רוחב גדול יותר מאפשר למודלים להתחרות בבקפ

Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning
במחקר חדש, נמצא כי טריינינג גרידי לייר-ווייס יכול להתחרות ביעילותו של טריינינג בקפ-אנד-פרופ. המחקר חוקר את השפעת רוחב המודל ועומקו על יעילות הטריינינג. נמצא כי במודלים רוחביים, טריינינג גרידי לייר-ווייס יכול להיות יעיל יותר. המחקר גם מצא כי המודלים הרוחביים הגרידי-לייר-ווייס יכולים ללמד דפוסי תיאור טובים יותר מאשר המודלים שנלמדו בטריינינג בקפ-אנד-פרופ.
תקציר מקורי באנגליתarXiv:2610.00753v1 Announce Type: cross Abstract: End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes achieve performance similar to backpropagation. However, the architectural conditions under which locally optimized networks, which avoid end-to-end backpropagation of error, can learn representations comparable to those learned through end-to-end training remain unclear. We aim to answer this question in the context of self-supervised learning, an important framework for large-scale pretraining in artificial intelligence. Here, we investigate how n
קרא במקור המקורי