יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ללא תוויות: קיצור זמן ראייה: צינורות תזיזה רק בתחום הטיה

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
במאמר זה, החוקרים מציגים פתרון לשאלה: האם ניתן לשפר את התפיסה של המודלים בזמן ראייה, כאשר הם פועלים בלי תוויות. הם מציגים פתרון שבו המודלים עובדים בלי תוויות, ומשפרים את התפיסה שלהם בזמן ראייה. הם מציגים תוצאות מעולות, ומציעים פתרון חדש לשאלה זו.
תקציר מקורי באנגליתarXiv:2609.18587v3 Announce Type: replace Abstract: Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the optimization space are severely restricted? We answer this question with label-free bias-only TTRL, which uses majority-vote pseudolabels as rewards and optimizes only ~100K bias parameters while keeping the pretrained backbone frozen. On MATH-500, our approach reaches 76.67% accuracy with Qwen2.5-7B, slightly exceeding our own labeled bias-steering reproduction while optimizing 76,000x fewer parameters than full-parameter TT
קרא במקור המקורי