כתבה
arXiv cs.CL ·
Vague2Detect: עיבוד פקודות רב-משמעיות בדטקציה של עולם פתוח
Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection
Vague2Detect מציע פתרון לבעיית פקודות רב-משמעיות בדטקציה של עולם פתוח. המערכת משתמשת ב-Sentence-BERT ו-GPT-3.5-turbo כדי לעבד פקודות ולזהות אובייקטים בתמונות.
תקציר מקורי באנגליתarXiv:2609.09949v1 Announce Type: cross Abstract: Real-world detectors must often interpret functional or ambiguous prompts, yet conventional models such as YOLO remain restricted to fixed class lists. Even open-vocabulary models like YOLO-World frequently misalign vague language with the intended objects. Building on our prior work Commonsense-Guided Open-World Object Detection Using LLMs and Visual-Semantic Matching, we address YOLO-World's limitations in grounding task-driven queries. We propose Vague2Detect, a hybrid pipeline in which a fine-tuned Sentence-BERT retrieves candidates from a structured household Knowledge Base (KB), and YOLO-World verifies their presence in the image. For prompts outside the KB, a large language model (GPT-3.5-turbo) generates candidate descriptions, dyna
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית