יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

עדיפות ריבוי מודלים על מודל בודד בתהליכי RLVR

Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It
חוקרים מצאו כי חלוקת תקציב RLVR למודלים רבים משפרת דיוק. שיטה זו, הנקראת 'adapter thicket', מאפשרת למודלים ללמוד מנתונים שונים ולשפר את הדיוק הכללי.
תקציר מקורי באנגליתarXiv:2610.00991v1 Announce Type: new Abstract: Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly $160$ completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ($1.5$B-$8$B). Yet on three of four models it leaves the majority vote below that of the untraine
קרא במקור המקורי