כתבה
arXiv cs.AI ·
Attribution in Scientific Literature: New Benchmark and Methods
תקציר מקורי באנגליתarXiv:2405.02228v4 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly generate citation-backed responses, yet citation hallucination remains a major challenge for trustworthy scientific information access. We introduce REASONS, a benchmark of 12,723 sentence-level citation instances spanning 12 arXiv subject categories, designed to evaluate scientific citation attribution under varying evidence conditions. We propose a dual-metric framework consisting of Abstention Rate (AR) and Hallucination Rate (HR) to characterize the trade-off between reliability and responsiveness. Using author-attribution and title-attribution tasks, we evaluate proprietary and open-source LLMs under zero-context, metadata-augmented, cascaded metadata-augmented prompting (CMP), retrieva
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית