יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למד עכשיו, שימוש בעתיד, אמון בעתיד: למידה טריוויאלית קודם לזמן הלמידה לאג'נטי LLM

Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents
אג'נטי LLM: למידה טריוויאלית קודם לזמן הלמידה. המאמר מציג פרקטיקה חדשה ללמידה בזמן הלמידה, שמאפשרת לאג'נטי LLM להתאים את עצמו לשינויים באירוע. הפרקטיקה, שנקראת StepLearn, משתמשת בטכניקה של למידה טריוויאלית כדי לעדכן את הידע של האג'נטי LLM בזמן הלמידה.
תקציר מקורי באנגליתarXiv:2609.35911v1 Announce Type: new Abstract: Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparametri
קרא במקור המקורי