יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בעיית הסימון בבדיקות הזייה

The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
חוקרים בדקו שיטות לזיהוי הזיות במודלים גדולים של שפה. הם מצאו חוסר התאמה בין שיטות סימון אוטומטיות ובין תגובות אנושיות.
תקציר מקורי באנגליתarXiv:2610.08026v1 Announce Type: new Abstract: In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false spe
קרא במקור המקורי