כתבה
arXiv cs.AI ·
הנחיה ללא תוויות: צינורת קיצור תפעול של רכיבת-למידה בשלב הבדיקה
Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
במאמר זה, המחברים מציגים טכניקה חדשה להנחיה ללא תוויות, המאפשרת צינורת קיצור תפעול של רכיבת-למידה בשלב הבדיקה. הם מציגים תוצאות מוצלחות על מספר משימות, כולל MATH-500, MathVista, AI2D, LogicVista ו-MMAU.
תקציר מקורי באנגליתarXiv:2609.18587v2 Announce Type: replace-cross Abstract: Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the optimization space are severely restricted? We answer this question with label-free bias-only TTRL, which uses majority-vote pseudolabels as rewards and optimizes only ~100K bias parameters while keeping the pretrained backbone frozen. On MATH-500, our approach reaches 76.67% accuracy with Qwen2.5-7B, slightly exceeding our own labeled bias-steering reproduction while optimizing 76,000x fewer parameters than full-parame
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית