יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מעבר מקומפילציה: בדיקת תאום טבעי של הצהרות-למידה על-פי-לשון- natural

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
בדיקת תאום טבעי של הצהרות-למידה על-פי-לשון- natural. ניתוח של שמונה מערכות, כולל GPT-5.2 ו-Gemini-2.5-Pro.
תקציר מקורי באנגליתarXiv:2606.31002v2 Announce Type: replace-cross Abstract: Lean verifies that a generated declaration is well typed, but not that it expresses the statement a user intended. We study two questions for autoformalization without canonical Lean targets: whether LLM judges can provide a usable proxy for human semantic review, and how much compilation overstates faithfulness across systems. Our criterion combines Lean compilation with strict semantic consensus between GPT-5.2 and Gemini-2.5-Pro. On an independently audited random sample, it agrees with human majority on 89.7\% of cases (Wilson 95\% CI: 82.1--94.3\%). Across eight systems evaluated on 400 graduate-level statements, every system has a nonzero compile--faithfulness gap, whose observed magnitude ranges from 3.0 to 29.0 percentage po
קרא במקור המקורי