כתבה
arXiv cs.AI ·
מתכון פתוח לזהב באולימפיאדה: אימון Nemotron למתמטיקה אולימפית
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
חוקרים אימנו את מודל Nemotron ליצירת הוכחות מתמטיות ברמת אולימפיאדה. המודל השיג 30 נקודות מתוך 42 באולימפיאדה 2026, ועבר את רף הזהב. החוקרים משחררים את המודלים המאומנים ואת קוד האימון.
תקציר מקורי באנגליתarXiv:2609.10712v1 Announce Type: new Abstract: We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained specialists - power an iterative search that generates, verifies, and refines candidate proofs; a separate high-compute stage then selects each final sub
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית