כתבה
arXiv cs.AI ·
SemVerBench: בניית סקאלה לבדיקת תחושת ה-LLM לפירוש סמנטיקה של תצורות גרסה
SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics
בניית סקאלה לבדיקת תחושת ה-LLM לפירוש סמנטיקה של תצורות גרסה. המאמר עוסק בבדיקת תחושת ה-LLM לפירוש סמנטיקה של תצורות גרסה, כולל שמות מודלים/כלים/חברות. המאמר עוסק בבדיקת תחושת ה-LLM לפירוש סמנטיקה של תצורות גרסה, כולל שמות מודלים/כלים/חברות.
תקציר מקורי באנגליתarXiv:2609.11180v2 Announce Type: replace Abstract: Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or >=2.0,<3, yet their grasp of version-constraint semantics has never been measured directly. We introduce SemVerBench, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo): 240 machine-checkable items with unique answers, built author-neutrally from four balanced sources (each ecosystem's official test suite plus three frontier LLM proposers) and labeled by a non-circular two-implementation oracle. Evaluating six frontier models, we find systematic, predictable per-mechanism blind spots: a partial-comparator carry rule (>1.2 means >=1.3.0) traps every model on Carg
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית