כתבה
arXiv cs.AI ·
GroundingPI: דגם יסודי לביסוס רב-תחומי כלפי מודלי תושייה פיזית
GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
GroundingPI היא דגם יסודי לביסוס רב-תחומי, המסוגלת לזהות ולמפות עצמים בעולם הפיזי. הדגם, המכונה GroundingPI, הוכשרה על ידי צוות מדעניות ומדעני טכנולוגיה מאוניברסיטת Stanford, והיא כבר נחשבת לאחת הדגמים הטובות בתחום.
תקציר מקורי באנגליתarXiv:2609.39601v1 Announce Type: cross Abstract: Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanni
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית