יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

האם סוכנים יודעים כשהם מצליחים?

Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
חוקרים שיטות למדידת ביטחון סוכנים מתוך ייצוגים פנימיים. הם מציגים שתי שיטות: Latent Trajectory Dynamics ו-Action Representation Probe, שמנבאות הצלחה מייצוגים שנוצרים בזמן קבלת החלטות. השיטות האלו עובדות עם מודלים כמו Qwen ו-DeepSeek.
תקציר מקורי באנגליתarXiv:2609.09448v1 Announce Type: new Abstract: As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions. In this paper, we investigate whether model's internal representations provide stronger signals of eventual task success in multi-turn agentic setups. We introduce two complementary methods: Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from representations formed at action de
קרא במקור המקורי