כתבה
arXiv cs.LG ·
מודלים לשפה לאחר אימון לביצועים זוכי מדליית זהב בתחרויות קוד
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
חוקרים פיתחו שיטה לשיפור ביצועי מודלי שפה בתחרויות קוד. הם השתמשו באימון מונחה ולמידת חיזוק כדי לשפר את הביצועים. המודל Nemotron-3-Nano-CC הגיע ל-291 נקודות לאחר אימון ו-468 עם אסטרטגיה חדשה.
תקציר מקורי באנגליתarXiv:2609.02849v2 Announce Type: replace Abstract: Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL). Using 22,000 curated problems, we train Nemotron-3-Nano-CC (30B-A3B) with SFT and RL and Nemotron-3-Ultra-CC (550B-A55B) with SFT alone. We further introduce GenCorrect, a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions. On IOI 2025, Nano-CC improves from 130 points to 291 after post-training and to 468 with GenCorrect, exce
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית