כתבה
arXiv cs.AI ·
שיפור פוליצי לקראת תקיפה מתקנה בשלב הבדיקה
Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation
CGPI מציע שיפור פוליצי שמשתמש בקרדיטים מוחזרים מהתקיפה לקראת תקיפה מתקנה. זה נעשה באופן זירו-שותף, כללי ומבוסס-קרדיטים. CGPI מציע שיפור פוליצי שמשתמש בקרדיטים מוחזרים מהתקיפה לקראת תקיפה מתקנה. זה נעשה באופן זירו-שותף, כללי ומבוסס-קרדיטים.
תקציר מקורי באנגליתarXiv:2609.37591v1 Announce Type: cross Abstract: Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience to correct such deviations. These signals, however, do not directly reveal whether an executed action supports instruction-guided progress toward the goal. Moreover, a plausible corrective signal does not guarantee a reliable policy update. The key challenge is thus twofold: identifying interactions that support goal-directed improvement and deter
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית