יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Efficiency Hallucination: תקיפה פורמלית ומדידה של תקיפה התנהגותית באופטימיזציה של קוד בעזרת LLM

Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization
LLMs נוטים לאופטימיזציה של קוד עם טענות בלתי-מוצדקות לשיפורי ביצועים. חברות: Claude, Gemini, GPT-5.
תקציר מקורי באנגליתarXiv:2609.14839v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into automated code optimization introduces a critical reliability risk we term the Efficiency Hallucination: an LLM's tendency to issue non-functional mutations with unsubstantiated performance claims on already-optimized code. This is driven by the Evaluation Trap, wherein binary benchmarks incentivize unnecessary modifications over safely abstaining. We present a validation framework using classification penalty methods, evaluated across 180 optimization runs on nine models (GPT, Claude, Gemini) using EffiBench. Under standard prompts, models exhibit a 100% over-edit rate on optimal code. Our guardrail raises correct abstention from 0% to to 44.4%, preserving a 100% edit rate on sub-optimal
קרא במקור המקורי