כתבה
arXiv cs.CL ·
חקירת גבולות ידע והתערבות מונחה על ידי דרישה לפיות קוד LLM לפענוח מערכות כוח
Knowledge boundary probing and demand-guided intervention for LLM-based power system code generation
מאמר זה עוסק בשיפור פעילות פיות קוד LLM לפענוח מערכות כוח, על ידי חקירת גבולות ידע והתערבות מונחה על ידי דרישה. המחברים מציגים את PowerCodeBench, סט של 2,000 תרגילים לבניית קוד, ופועלים לשפר את דיוק הפיות על ידי שימוש בטכניקות של חקירת גבולות ידע והתערבות מונחה על ידי דרישה.
תקציר מקורי באנגליתarXiv:2605.31478v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can turn grid-analysis requests into executable programs for power-system simulation, but utilities and research laboratories often require on-premise deployment. In this setting, first-pass failures frequently arise at an API-knowledge boundary, through hallucinated functions, misused parameters, and mishandled result tables. We present PowerCodeBench, a parameterised benchmark generator released as a frozen 2,000-task suite for pandapower, and a deployment-time workflow that requires no weight updates. Documentation-driven L0-L3 probes produce per-model API profiles for diagnosis, model comparison, documentation allocation, and backend calibration. A query-side demand estimator selects layered API evid
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית