יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

שיפור זרימות תבוניות במודלים רב-מודאליים

Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models
FLOWMEM הוא מודל רב-מודאלי ששומר ומשפר זרימות תבוניות לצורך בקרה רובוטית. הוא משתמש ברכיבים קודמים כדי ליצור נתיבי תבונית חדשים. ניסויים הראו שיפור של 1.7-4.1% בהצלחה.
תקציר מקורי באנגליתarXiv:2610.12090v1 Announce Type: new Abstract: Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions. While existing methods learn to generate or refine such states for each policy query, they discard successful reasoning after execution and therefore reconstruct similar computation from scratch. We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience. Rather than appending a fixed retrieved context, FLOWMEM dynamically retrieves and recomposes compatible latent fragments as the embodied context evolves, forming a reasoning route that follows the temporal structure and progress of s
קרא במקור המקורי