כתבה
arXiv cs.AI ·
Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics
תקציר מקורי באנגליתarXiv:2609.37907v2 Announce Type: replace Abstract: Video games offer scalable environments for studying perception and control in embodied agents. Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on $\sim$1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית