כתבה
arXiv cs.AI ·
גיבוש מודלים LLM עם ראיות רועשות
How Well Do LLMs Reason with Noisy Evidence? An Active Visual Reasoning Benchmark
חוקרים פיתחו בנך' VisualNoiseQA לבדיקת יכולתם של מודלים LLM לגיבוש עם ראיות רועשות. הבנך' מאפשר למודלים לשאול שאלות ולקבל תשובות עם אי-ודאות. הניסויים הראו כי המודלים יכולים לנצל אותות אי-ודאות לגיבוש עמיד.
תקציר מקורי באנגליתarXiv:2610.07751v1 Announce Type: new Abstract: Real-world reasoning rarely reduces to static question answering: agents must actively gather information from tools and sensors that are often noisy and unreliable. Yet most existing active reasoning benchmarks assume that environmental feedback is trustworthy, or introduce noise without exposing an explicit, calibrated uncertainty signal, leaving open how LLMs should reason when the evidence itself is uncertain. We introduce VisualNoiseQA, a novel benchmark for active reasoning under noisy visual feedback. A text-only LLM must solve VQA problems by iteratively querying a fixed, off-the-shelf VLM treated as a stochastic visual sensor. For each query, we draw multiple samples and expose an empirical uncertainty signal via self-consistency, en
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית