כתבה
arXiv cs.AI ·
EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
תקציר מקורי באנגליתarXiv:2610.03710v1 Announce Type: cross Abstract: Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית