כתבה
arXiv cs.CL ·
איתור שגיאות עובדתיות בטקסטים שנוצרו על ידי צ'אטבוטים
Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection
חוקרים פיתחו שיטה לאיתור שגיאות עובדתיות בטקסטים שנוצרו על ידי צ'אטבוטים. השיטה כוללת שילוב של אנוטציה ראשונית, גילוי מועמדים באמצעות LLM ואדג'ודיקציה רפואית ובדיקת עובדות. התוצאות מראות כי שיטה זו יכולה לשפר את הדיוק באיתור שגיאות.
תקציר מקורי באנגליתarXiv:2609.03953v1 Announce Type: new Abstract: Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass, single-annotator labeling problem. In long-form chatbot responses, factual errors can be subtle and embedded within mostly correct text. We develop a multi-perspective annotation study of medically relevant chatbot responses, combining first-pass annotation, LLM-as-a-Judge (LaJ) candidate discovery, and two forms of adjudication: medical-expert and evidence-based fact-checking. First-pass annotators frequently miss factual errors later validated by adjudicators. LaJ improves candidate discovery, but is insufficient on i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית