כתבה
arXiv cs.AI ·
MAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
תקציר מקורי באנגליתarXiv:2608.16697v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However, enabling VLAs to self-evaluate their action generation reliability without external supervision remains a major challenge. Existing methods either rely on expert annotations or estimate uncertainty only from output statistics, largely ignoring internal signals. In this work, we observe that internal visual modality entropy exhibits consistent distinctions between successful and failed tasks across heterogeneous VLAs. Although VLAs' architectures differ in their action generation, we show that they share a common latent action generation abstraction evolving unde
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית