כתבה
arXiv cs.AI ·
כללים לכלים: בדיקות מבצעות לסוכנים LLM בחישוב מדעי
Rules to Tools: Executable Checks for LLM Agents in Scientific Computing
כללים לכלים (R2T) מספקים בדיקות מבצעות לדרישות מדעיות ציבוריות. הכלי מאפשר לסוכנים LLM לבצע בדיקות ולתקן תוכניות. המחקר מראה תוצאות טובות עם השימוש בבדיקות המבצעות.
תקציר מקורי באנגליתarXiv:2610.00313v1 Announce Type: new Abstract: Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requirements. Matched SciCode repair groups share written checks, starting programs, model, and budgets; the tool group receives a callable implementation. Across two task-ID cohorts, complete repair is 26/30 with text and 29/30 with the prepared checks. Three task IDs favor tools, one favors text, and eleven tie. The eight-ID cohort scores 13/16 versus 15/16, with a task-cluster bootstrap 95% interval of [-12.5, 43.75] percentage points for the difference. The larger shared-definition SciCode cohort ties at 13/24 per group.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית