יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

הזיות במודלים רב-מודאליים

Hallucination in Multimodal Foundation Models: A Survey on Causes, Corrections, and Evaluations
מודלים רב-מודאליים מאופיינים ביכולת לעבד מידע ממקורות שונים. הזיות במודלים אלו עלולות לפגוע באמינות וביעילות. סקירה זו מציגה ניתוח מעמיק של סיבות, תיקונים ושיטות להערכת הזיות.
תקציר מקורי באנגליתarXiv:2410.15359v2 Announce Type: replace Abstract: Multimodal Foundation Models represent a significant leap in artificial intelligence. Among them, Large Vision-Language Models (LVLMs) serve as the typical representative of these foundation models, which integrate visual modality directly into Large Language Models (LLMs). They have demonstrated strong capabilities in information processing and generation. However, the existence of hallucinations has limited the potential and practical effectiveness of LVLM in various fields. Although lots of work has been devoted to hallucination mitigation and correction, there are few reviews to summarize them. To address this gap, this survey provides a systematic review of the hallucination landscape in LVLMs. We categorize the causes related to mod
קרא במקור המקורי