כתבה
arXiv cs.AI ·
SCOPE: אימות משפטים בעזרת מודל שפה
SCOPE: Certified Theorem Proving with a Language Model as the Policy Planner
SCOPE הוא כלי לאימות משפטים מתמטיים בעזרת מודל שפה. הוא משתמש במנוע סימבולי לביצוע חישובים ובמהדר ליצירת הוכחות. בניסויים, SCOPE אימת 191 מתוך 218 משפטים עם גב למודל בגודל 135M. הוא עובד טוב יותר מDeepSeek-Prover-V1.5-RL ו-DeepSeek-Prover-V2-7B.
תקציר מקורי באנגליתarXiv:2610.08319v1 Announce Type: new Abstract: In proof assistants such as Lean, a generated proof must pass machine compilation checks, so evaluation needs no human scoring. Direct generation fails on multi-step numeric propositions: a proof is valid only if every content integer is correct, so the pass rate is bounded by the k-th power of the per-integer accuracy. Controlled corruption across 2,617 reference proofs confirms this power law. SCOPE (State-Conditioned Operator Planning and Execution) enforces the natural division of labor: the model plans over an operator vocabulary, a symbolic engine executes the numerics, and a compiler renders the proof. On a 218-problem suite it certifies 191/218 (87.6%) with a 135M backbone; the 7B DeepSeek-Prover-V1.5-RL certifies 18/218 at 27.5 times
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית