כתבה
arXiv cs.AI ·
CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?
תקציר מקורי באנגליתarXiv:2609.37216v1 Announce Type: cross Abstract: Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJudgeBench, a benchmark of 1199 instances constructed from real pull requests and expert-verified perturbations, covering both trustworthy and plausible
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית