כתבה
arXiv cs.LG ·
Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
תקציר מקורי באנגליתarXiv:2606.31413v2 Announce Type: replace-cross Abstract: Composing independently trained LoRA adapters into a single large language model is useful for multi-domain adaptation, especially when the original training data cannot be shared. A common approach is to use MoE-style routing over LoRA experts, but for frozen pretrained adapters, soft weighted combinations can change the unit-scale additive update under which each LoRA module was originally trained. We propose \textbf{Hard-Routed MoR-LoRA}, a two-stage framework for composing frozen reasoning LoRA experts through unit-scale hard selection. First, domain-specific LoRA adapters are trained independently using reinforcement learning from verifiable feedback to obtain reasoning experts. Then, all experts are frozen, reasoning traces ar
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית