כתבה
arXiv cs.AI ·
ViSR-KGC: שילוב מודלים ויז'ואליים-לשוניים להשלמת גרפים ידע רב-מודאליים
ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion
ViSR-KGC הוא גישה חדשה להשלמת גרפים ידע רב-מודאליים, המשלבת למידת ייצוגים, ניתוח ראיות רב-מודאליות וידע מוקדם. היא מאפשרת למודלים ויז'ואליים-לשוניים לנתח גרפים ולהשלים ישויות חסרות.
תקציר מקורי באנגליתarXiv:2608.05833v3 Announce Type: replace Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images. Traditional representation learning approaches follow the embedding-based paradigm and may struggle when relation-specific evidence is limited. Meanwhile, LLM-based reasoning methods typically linearize graph structures into textual prompts, which obscures structural topology and neglects vital visual information. While vision-language models (VLMs) excel at multimodal reasoning, they cannot natively interpret structured graph topology, particularly when it comes to knowledge graphs w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית