כתבה
arXiv cs.LG ·
בדיקה סגורה של סוכני LLM לפיתוח תוכנה
Closed-loop evaluation of LLM agents for embedded software development
בדיקה סגורה של סוכני LLM לפיתוח תוכנה. המחקר מציג בנק אבות טיפול סגור של סוכני LLM לפיתוח תוכנה נטענת. הבנק כולל חמש תפקודיות של שליטה נשלטת וארבע סקנריות של חזרה: ייצור תקין, עקיפה ממשית, חזרה CI וחזרה אורקל.
תקציר מקורי באנגליתarXiv:2610.11447v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as coding agents that edit files, run builds and tests, inspect execution results, and repair software iteratively. Embedded firmware is a demanding target because correctness depends on closed-loop behavior under sensing, timing, and safety constraints, not only on static source quality. Yet embedded-agent evaluation remains limited and often emphasizes one-shot synthesis or offline correctness. We present a benchmark for closed-loop evaluation of embedded coding agents. Each task provides a plain-text engineering description, constrained workspace, and visible build-and-runtime surface. The agent must translate requirements into implementation and self-verification steps, then iterate
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית