יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

כאשר ציפורניות UQ סופרוויזד השיפוטים משפרות את זיהוי ההלוצינציות של LLM?

When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study
חידוש בשיפור זיהוי הלוצינציות של LLMs: חקירה על ציפורניות UQ סופרוויזד.
תקציר מקורי באנגליתarXiv:2608.24492v2 Announce Type: replace Abstract: Uncertainty quantification (UQ) methods are widely used for hallucination detection in large language models (LLMs) in closed-book settings where ground-truth evidence is unavailable at inference time. Prior work has proposed combining UQ signals via learned ensembles, but empirical investigations into the robustness of these ensembles are limited. We study a supervised ensembling framework that trains a classifier over heterogeneous UQ-based scorer outputs on a small, domain-specific dataset of labeled LLM responses, then applies it to out-of-sample hallucination classification without retrieval, tools, or reference documents. Across four LLMs, nine datasets, and three generation regimes (short-form QA, long-form generation, and code gen
קרא במקור המקורי