יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הערכת למידה עמוקה אישית עם התערבויות זמניות

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
חוקרים הציעו פרוטוקול חדש להערכת סוכנים אישיים של למידה עמוקה. הפרוטוקול בוחן את היכולת של הסוכנים להתמודד עם התערבויות זמניות ולשמור על יציבות. המחקר מציע כי פרוטוקול זה יכול לשפר את הבטיחות והיעילות של סוכנים אישיים.
תקציר מקורי באנגליתarXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall or forgetting, and safety benchmarks test static policy compliance. We argue that personal-agent evaluation requires a different protocol: replaying the same temporal intervention across different persistent user-conditioned states and measuring how failures propagate across agent components. We formalize this requirement as four conditions: explicit temporal intervention, persistent state across the intervention, induced cross-dimensional effects, and variation in user-condit
קרא במקור המקורי