כתבה
arXiv cs.AI ·
SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization
תקציר מקורי באנגליתarXiv:2608.29270v3 Announce Type: replace-cross Abstract: Autoformalization translates informal mathematical theorems into code for proof assistants such as Lean. A central challenge is that current evaluation metrics can accept type-correct but misaligned statements or reject correct statements written in a different formulation. Inspired by Pass@$k$, we propose SA-Pass (*Semantic Alignment Pass*), which tests formal statements using auxiliary statements called *shadows* that characterize the intended statement. A generated statement receives full credit only when it compiles, implies each shadow (forward check), and is implied by their conjunction (backward check). We instantiate SA-Pass in ShadowBench, a Lean 4 full autoformalization benchmark of 178 postgraduate- to research-level prob
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית