כתבה
arXiv cs.CL ·
CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models
תקציר מקורי באנגליתarXiv:2606.19788v2 Announce Type: replace-cross Abstract: We present CombEval, a dynamic benchmark for evaluating combinatorial counting in large language models. CombEval represents each problem as a typed Cofola specification over entities, combinatorial objects, object dependencies, and constraints, enabling controlled generation of natural-language counting problems with exact solver-verified answers. Unlike static collections, CombEval supports systematic variation of object type, entity scale, constraint count, and reasoning depth. We evaluate 11 LLMs under direct and code-augmented settings and find that models remain brittle on ordered objects, indistinguishable elements, relatively positional constraints, and nested object dependencies. Error analysis further identifies failures i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית