יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

CLAMP: קונסטריינט דיקודינג לתכנון עם תכונות ראייה-שפה

CLAMP: Constrained Decoding for Vision-Language Embodied Planning
CLAMP היא תשתית קונסטריינט-גראונדינג שמגבילה את הדיקוד של תכנון עם תכונות ראייה-שפה. היא משתמשת בהנחיות סימבוליות ובעזרת רשת חשבוני מצב (HMM) כדי למנוע קונדיטורים לא תקינים. CLAMP נבחנה ב-VLABench, SafeAgentBench ו-TaPA.
תקציר מקורי באנגליתarXiv:2609.08602v1 Announce Type: new Abstract: Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequences. However, fluent plans are not always executable. A VLM may refer to objects that are not visually observed, select actions whose required affordances are unavailable, or violate syntax and action constraints. We introduce CLAMP, a multimodal constraint-grounding framework that turns scene evidence into decoding-time constraints for a frozen VLM planner. CLAMP uses the initial observation to restrict object references to those supported by the scene, while a provided symbolic action model specifies state transitions and goals. During decoding, hard masks eliminate invalid next-token candidat
קרא במקור המקורי