כתבה
arXiv cs.AI ·
ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation
תקציר מקורי באנגליתarXiv:2607.00711v2 Announce Type: cross Abstract: Large Language Models have emerged as programming assistants. However, the efficacy of code generation is constrained by the quality of input requirements, which are frequently ambiguous, incomplete, or underspecified. While LLMs excel at one-shot code synthesis, their ability to proactively clarify intent remains underexplored, as a critical trait for robust software engineering. Existing benchmarks largely overlook this interactive bottleneck, assuming perfectly specified prompts that do not reflect the iterative nature of requirement elicitation. To bridge this gap, we introduce ClarifyCodeBench, a novel interactive benchmark for evaluating LLMs' capability in resolving requirement ambiguity. Constructed from real-world programming tasks
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית