כתבה
arXiv cs.CL ·
ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
תקציר מקורי באנגליתarXiv:2610.12403v1 Announce Type: cross Abstract: Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure. Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored. We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents. Retrieved skills guide both inference and reward shaping, while successful trajectories are distilled back into the library, forming
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית