יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מודלי שפה לאחר הכשרה: הצלחה זהב בתחרויות תכנות

Post-Training Language Models for Gold-Medal Performance in Coding Competitions
מודלי שפה שהוכשרו לאחר הכשרה הצליחו לשחק על גבול הזהב בתחרויות תכנות. המחקר מציג פיפליין של הכשרה של מודלי שפה גדולים, כולל סינתטי, הכשרה סופרוויזד, ולמידת רפלקסיה. המחקר מציג גם את GenCorrect, סטרטגיה של חיפוש ושיפור בזמן ריצה. המחקר מציג תוצאות של 535.4 נקודות, על פי IOI 2026, ומציע פיתוח של מערכת Ultra-CC.
תקציר מקורי באנגליתarXiv:2609.02849v2 Announce Type: replace-cross Abstract: Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL). Using 22,000 curated problems, we train Nemotron-3-Nano-CC (30B-A3B) with SFT and RL and Nemotron-3-Ultra-CC (550B-A55B) with SFT alone. We further introduce GenCorrect, a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions. On IOI 2025, Nano-CC improves from 130 points to 291 after post-training and to 468 with GenCorrect
קרא במקור המקורי