יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מתי תיקון הופך לתיקון? ניתוח מניפולטיבי של התערבויות פנימיות ב-LLMs

When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs
ניתוח מניפולטיבי של התערבויות פנימיות ב-LLMs
תקציר מקורי באנגליתarXiv:2609.36138v1 Announce Type: new Abstract: Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; across three sealed evaluations, none of 59 budget-matc
קרא במקור המקורי