כתבה
arXiv cs.AI ·
קוד שעובד, אך אינו עובד: מדידת חוזרינות של סביבות באי-גנרציה של קוד
Code That Works, Environments That Don't: Measuring Environment Reproducibility in AI-Generated Software
במאמר זה, נבחן את יכולתן של סוגיות קוד לזהות תלותיות ראשוניות, ולהשוות את תוצאותיהן לאלו של סוגיות קוד שנכתבו ביד אנוש.
תקציר מקורי באנגליתarXiv:2610.00425v1 Announce Type: cross Abstract: Code generation has emerged as a central capability of large language models, with coding agents now able to produce functionally correct software projects from natural language prompts. However, functional correctness alone does not capture a critical dimension of generation quality: environment specification, defined as the accurate identification of the dependencies required to execute generated code, is equally critical. We develop an agent protocol for environment specification and introduce a three-layer framework comprising declared, runtime-installed, and necessary-and-sufficient dependencies to systematically assess coding agents for environment specification. Using this protocol, we evaluate the extent to which coding agents syste
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית