כתבה
arXiv cs.AI ·
CulturalMenuBench: חקירת הפער בין ידע ליישום בתחום ההסבר המולטימודלי של המזון
CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
CulturalMenuBench - תקן חדש שחושף פער בין ידע ליישום בתחום ההסבר המולטימודלי של המזון. התקן כולל 4,870 פריטים ב-10 שפות ו-18 אזורים, ובודק את יכולת המודלים להסביר את המזון ואת הטכניקות הבישוליות. התקן חשף פער בין ידע ליישום, ומציע תרגילים חדשים להסבר המולטימודלי של המזון.
תקציר מקורי באנגליתarXiv:2609.03526v1 Announce Type: new Abstract: Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית