כתבה
arXiv cs.LG ·
אמון בכיוון, חיפוש בצעד: תיאוריות ראשוניות ואפסיות להטמעת LLM
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
במאמר זה, החוקרים מציגים תיאוריה חדשה לבחירת צעדים בהטמעת LLM. התיאוריה, הנקראת ZFO, משתמשת באופטימיזציה ראשונית כדי לבחור בכיוון ובוחרת בצעד על ידי חיפוש זריז. התיאוריה נותנת הבטחות תאורטיות לכך שהיא תקפה ומשפרת את הביצועים של LLM.
תקציר מקורי באנגליתarXiv:2610.02190v1 Announce Type: new Abstract: Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded searc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית