כתבה
arXiv cs.LG ·
בחינה של סוכני דגלי שפה עבור תהליכי תכנות ניידים: מחקר תכנותי רפליקציה של עבודות ניידות
Evaluating Local Language Model Agents for Reproducible Data Engineering: An Empirical Software Engineering Study of Mobility Workflows
במאמר זה, נבחן סוכני דגלי שפה עבור תהליכי תכנות ניידים. נבחן את יכולתם של סוכנים אלה לייצר תוצאות נאמנות ורפליקציה, ולבדוק את השפעת תנאי עבודה סגורים. נמצא כי סוכני דגלי שפה יכולים לתמוך בתהליכי תכנות ניידים, אך יציבותם תלויה ביכולת המודל, באפשרות לבדוק תוצאות ובאפשרות לבדוק תוצאות.
תקציר מקורי באנגליתarXiv:2610.11482v1 Announce Type: cross Abstract: Context: Large language model (LLM) agents are increasingly used as software and data-engineering assistants, yet evidence about locally deployable open-weight agents remains limited. Existing evaluations often emphasize textual responses or isolated code generation rather than the validity of complete engineering artifacts. Objectives: We evaluate whether local LLM agents can produce correct and reproducible data-engineering artifacts, quantify the effect of a closed-loop workspace condition, and examine trade-offs in model scale, architecture, quantization, runtime, tool use, and failure. Methods: We introduce a benchmark of fifteen mobility-workflow tasks covering data discovery, connectors, transport-feed processing, semantic enrichment
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית