כתבה
arXiv cs.AI ·
SyntaxBench: תשתית רפואית סטטיסטית לשיפוט סינטקסי במודלי שפה גדולים
SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models
תשתית חדשה לשיפוט סינטקסי במודלי שפה גדולים. התשתית כוללת 6 משימות, כולל ספירת תווים, חיפוש תווים, זיהוי פלינדרום, מרחק עריכה ובחירת המחרוזת הארוכה ביותר. התשתית נבדקה על 8 מודלי שפה פתוחים, והתוצאות הראו שהמודלים חסרים בשיפוט סינטקסי.
תקציר מקורי באנגליתarXiv:2610.03329v1 Announce Type: cross Abstract: Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts. We evaluate eight
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית