יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הגברת המס בפוסט-אימון

Sharpening Tax in Post-Training
בפוסט-אימון, המודלים נוטים להגביר את התנהגויות הקיימות, אך עלולים להגביל את כיסוי הפתרונות. המחקר מציג חידוש חדשני, PTGS, שמאפשר תפקוד טוב יותר של המודלים בסביבות רב-פסים.
תקציר מקורי באנגליתarXiv:2610.01509v1 Announce Type: cross Abstract: An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying me
קרא במקור המקורי