יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אמינות במודלים גדולים של שפה

Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective
חוקרים הציגו את PFaithBench, בנק אמת לאמינות במודלים גדולים של שפה. המחקר בוחן 39 מודלים ומוצא כי רובם נוטים לתת תשובות לא מדויקות. הקוד והנתונים זמינים ב-GitHub.
תקציר מקורי באנגליתarXiv:2610.07894v1 Announce Type: new Abstract: Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically assess each question-context instance in isolation; however, such instance-level evaluation fails to capture a fundamental requirement of faithful behavior: the ability to adapt model responses to changes in available contexts. In particular, a model should provide correct answers when sufficient evidence is present and abstain when it is not. In this work, we propose a Pairwise Faithfulness Benchmark (PFaithBench) that evaluates whether a model can switch between answering and abstaining for the same question under s
קרא במקור המקורי