יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תיקון FOLIO ו-MALLS: אפיון מאומת ופלטפורמה של LLM להקצאת תיקון אנושי

Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling
אפיון מאומת של נתוני FOLIO ו-MALLS, ופלטפורמה של LLM לתיקון אנושי. נמצאו תקלות ב-42.5% ו-42% מהנתונים, והתיקון גרם לשיפור בדיוק של +11 ל-+23%.
תקציר מקורי באנגליתarXiv:2606.02837v2 Announce Type: replace Abstract: Accurate translation from Natural Language to First-Order Logic (NL-to-FOL) underpins neurosymbolic AI systems and Natural Language Inference (NLI), making the quality of NL-to-FOL benchmarks essential---yet these datasets have never been rigorously audited. Our first contribution is to present a systematic human inspection of the validation split of \textsf{FOLIO} and a subset of \textsf{MALLS} test instances, finding that approximately 42.5\% and 42\% of entries, respectively, contain incorrect FOL formalizations (i.e., ground truth labels), with additional rates of ambiguous NL sentences (17.8\% and 51\%) and incorrect NLI labels in \textsf{FOLIO} (8.4\%). Our second contribution is to develop and release corrected ground truths for su
קרא במקור המקורי