יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חידוד מס באימון לאחר הכשרה

Sharpening Tax in Post-Training
חוקרים גילו שאימון לאחר הכשרה של מודלים גדולים של שפה (LLM) יכול לשפר דיוק אך על חשבון כיסוי פתרונות. הם הציגו את מדד 'Sharpening Tax' למדידת העלות. המחקר בדק 14 זוגות של מודלים והראה תוצאות משמעותיות.
תקציר מקורי באנגליתarXiv:2610.01509v1 Announce Type: new Abstract: An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mech
קרא במקור המקורי