יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

באיזה נקודות חוקרים צירופי סוכנים תפיסה של עמידה לעומת התקדמות?

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
צירופי סוכנים טועים בין עמידה להתקדמות עקב תפיסה עצמית. חוקרים חקרו את הבעיה ומצאו שהסיבה היא תלות במקור האמת.
תקציר מקורי באנגליתarXiv:2607.25152v2 Announce Type: replace Abstract: Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, ye
קרא במקור המקורי