יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת מנת התנהגות לאימון LLM

Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents
BQ-LoRA הוא כלי לאימון סוכנים LLM. הוא משתמש במנת התנהגות מקומית כדי לארגן עדכוני מסלול. המערכת כוללת שני מודולים: איזון מנת התנהגות ודחיסת שימור החלטות. BQ-LoRA מושווה לשיטות אימון אחרות בניסויים.
תקציר מקורי באנגליתarXiv:2609.12896v1 Announce Type: new Abstract: LLM-based agents rely on heterogeneous interaction capabilities to accomplish complex tasks. Existing approaches often distribute these capabilities across multiple LoRA adapters, which increases adapter storage requirements and introduces routing overhead during inference. A single LoRA avoids this overhead, but learning from diverse agent trajectories under a fixed rank budget presents two challenges. First, trajectories with different interaction traces and parameter gradients can induce equivalent changes in decision distributions, causing repeated updates to overemphasize redundant behavioral changes. Second, an aggregated update may exceed the rank budget of the adapter, and approximating it in weight space can distort the decision chan
קרא במקור המקורי