יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מבחינה בברירה: לימוד לשלוט באגנטי LLM

From Uncertainty to Action: Learning to Steer LLM Agents
במאמר זה, המחברים חוקרים את האפשרות לשלוט באגנטי LLM באמצעות ביטול וערך של שליטה. הם מציגים את VoS, מנוע שליטה שלימד מערך סטפ-בי-סטפ (SOT) את ערך השליטה בכל שלב. VoS משפר את הביצועים של האגנטים ב-12 סיטואציות שונות.
תקציר מקורי באנגליתarXiv:2610.09115v1 Announce Type: cross Abstract: Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and run each continuation to completion. The resulting stepwise outcome table (SOT) holds about 82,000 counterfactual continuations of 1,864 trajectories from three benchmarks and two agents. It shows that uncertainty can identify failing trajectories, but that no single signal reliably locates the step at which steering helps. We therefore propose VoS (Value of Steering), a trajectory-level monitor, offline or online, that learns
קרא במקור המקורי