כתבה
arXiv cs.CL ·
סנטנס-לווייל תרגול-חופשי כמדד מחקר-חינם של תוכן לא-מקובל
Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers
מדד חדש ללא-תרגול לזיהוי תוכן לא-מקובל בתשובות RAG
תקציר מקורי באנגליתarXiv:2607.04223v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) assistants summarize records in clinical and legal work, where one unsupported sentence can mislead a reader. The contrast between an output's likelihood with and without its source is an established faithfulness score for whole summaries and answers, but it has not been measured as a detector of the individual unsupported sentence in multi-passage RAG answers, against trained verifiers, or for its cost. We implement it as a training-free detector that re-scores a fixed answer under the full context, no context, and each chunk removed, and returns the chunk whose removal lowers a sentence's likelihood most as a candidate supporting passage. We evaluate it on RAGTruth, TofuEval, and RAGBench with six sc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית