יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

LoGRA: פיתוח LLM עם רכיבי רגרדיאנטים נמוכי דרגה

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches
LoGRA מציגה שיטה להפחתת זיכרון באימון LLM עם רכיבי רגרדיאנטים נמוכי דרגה. השיטה תומכת בעדכוני דגל וסינכרון פוליצי באופן יעיל.
תקציר מקורי באנגליתarXiv:2610.06647v2 Announce Type: replace Abstract: Reinforcement learning has greatly advanced the capabilities of large language models, but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predicted-KL step control, which estimates policy changes before applying each update and adjusts its magnitude accordingly. With all techniques combined, LoGRA reduces average training memory usage by up to 45.7% across reasoning tasks without compromising performan
קרא במקור המקורי