יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

צעד אחד, נהיגה אחת: הפחתת התערבות של מסדרי גבוה בלמידת רפלקסיה מרובה-תחומים על ידי שליטה מעבר לצעד

One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control
הפחתת התערבות של מסדרי גבוה בלמידת רפלקסיה מרובה-תחומים על ידי שליטה מעבר לצעד. המחברים מציגים את OSOL, שמטרתו למנוע תערבות זו על ידי ניהול תגובה עקיפה. המחקר מדגים את יעילות OSOL בשלושה תחומים שונים: רפלקסיה, תגובה והתאמה.
תקציר מקורי באנגליתarXiv:2609.06469v1 Announce Type: new Abstract: Reinforcement learning (RL) across multiple domains can broaden the reasoning capabilities of large language models (LLMs), yet joint training often degrades individual-domain performance and can destabilize optimization. Existing work typically diagnoses such interference from a single-step view using first-order gradient alignment or curvature-based proxies. We show that this view can miss a critical form of sequential interference: same-point domain gradients may remain nearly orthogonal even when consecutive realized updates partially reverse one another in output space. We further show that consecutive token log-probability footprints recover this interaction directly from adjacent checkpoints as a local second-order interaction in outpu
קרא במקור המקורי