כתבה
arXiv cs.AI ·
VICO: סביבות חזותיות שמתפתחות במקביל למודלי תקשורת ראייה-לשון
VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
VICO: סביבות חזותיות שמתפתחות במקביל למודלי תקשורת ראייה-לשון. המאמר מציג פרקטיקה חדשנית לאימון מודלי VLM שמשתמשת בשינוי סביבות האימון במקביל לאימון המודל.
תקציר מקורי באנגליתarXiv:2610.10782v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית