יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

4MT-VLM: עד כמה גסה המפה הקוגניטיבית של VLM?

4MT-VLM: How Coarse Is a VLMs Cognitive Map?
4MT-VLM הוא מאגר נתונים של נוף מושקע, המבדק את היכולת של מודלים כמו Gemini ו-GPT לזהות מקומות מנקודות מבט שונות. התוצאות מראות כי המודלים מתקשים לשמור על מפה קוגניטיבית יציבה כאשר הנקודת המבט משתנה.
תקציר מקורי באנגליתarXiv:2609.39238v1 Announce Type: new Abstract: An agent that moves must recognise a place from a viewpoint it has never seen. We introduce 4MT-VLM, a dataset of procedurally generated landscapes, each rendered across five stimulus modes that remove appearance cues while holding layout fixed: shape and colour, shape only, colour only, bare terrain peaks with no objects, and a valley viewpoint that puts the peaks on the horizon. The last condition is commonly used in clinics to probe hippocampal function in human patients. We test this benchmark across sixteen different open and closed-source models and report 4AFC performance, a measure which is also used to grade human participants. We observe that models identify a place from the studied viewpoint but lose it once the camera moves, dropp
קרא במקור המקורי