כתבה
arXiv cs.AI ·
IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective
תקציר מקורי באנגליתarXiv:2609.15387v1 Announce Type: cross Abstract: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the application, yet incomplete exploration can cause them to miss implemented functionality and confound application defects with agent execution failures. To address these limitations, we propose IWC-Bench, an interactive benchmark for evaluating web application generation from a software testing perspective. IWC-Bench instruments each generated application and uses code coverage to guide an agent in exploring its functionality through
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית