כתבה
arXiv cs.AI ·
תפיסה מושכלת בחזות: יסודות, התקדמויות אחרונות וכיוונים עתידיים
Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions
מאמר זה מציג תפיסה מושכלת בחזות, המשלבת נתוני חזות וידע תקופתי. זהו צעד חשוב לשפר את ההבנה של המודלים האינטליגנטיים, ולאפשר להם לשתף פעולה עם בני אדם ועם הסביבה. המאמר כולל דיווח על התקדמויות האחרונות בתחום זה, ומציע כיוונים עתידיים לפיתוח מערכות חזותיות מושכלות.
תקציר מקורי באנגליתarXiv:2609.05257v1 Announce Type: new Abstract: Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning models but also enhances their ability to interact meaningfully with humans and the environment. Unlike CNN-based conventional vision models, which are designed to identify objects within a specific image, incorporating commonsense knowledge enables models to interpret scenes in a more holistic manner, thereby improving their spatial ability to reason about relationships among objects and actions. This integration not only enhances object recognition but also facilitates a deeper understanding of the contextual factors, ultim
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית