כתבה
arXiv cs.CL ·
OmniHallu: איתור אחיד של הזיות במודלים רב-מודאליים
OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models
OmniHallu הוא כלי לאיתור הזיות במודלים רב-מודאליים. הוא פותר את הבעיה של הזיות בפלטים של מודלים אלו. OmniHallu-Bench הוא בנק אימות עם 10,000 דוגמאות.
תקציר מקורי באנגליתarXiv:2609.11244v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse tasks, they suffer from hallucinations where generated outputs contradict or misrepresent input semantics. Existing research typically addresses hallucination detection within a single modality or task type, limiting generalizability. We introduce OmniHallu, a unified hallucination detection framework spanning both comprehension and generation tasks across image, video, and audio modalities. We contribute OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations covering six cross-modal tasks: image-to-text (I2T), video-to-text (V2T), audio-to-text (A2T), text-to-image (T2I), text-to-video (T2V), and text-to-audio (T2A). Our mul
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית