כתבה
arXiv cs.AI ·
למד עכשיו, שימוש בעתיד, אמון בעתיד: למידה טריוויאלית קודם-פרקוויאלית לאג'נטי LLM
Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents
אג'נטי LLM רוכשים ניסיון ומתאימים לשימוש בעתיד. המאמר מציג פרקטיקה ללמידה טריוויאלית קודם-פרקוויאלית שמאפשרת לאג'נטי להתאים לשימוש בעתיד. הפרקטיקה, שנקראת StepLearn, מספקת השגה ואמון בעתיד. המאמר מציג תוצאות של ניסויים שהראו תוצאות טובות יותר מאשר פרקטיקות אחרות.
תקציר מקורי באנגליתarXiv:2609.35911v1 Announce Type: cross Abstract: Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparamet
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית