יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ייצור רפרנסים ביו-רפואיים נותר לא אמין

Biomedical Reference Generation Remains Unreliable across 26 Large Language Models
ייצור רפרנסים ביו-רפואיים על ידי 26 מודלים שונים, כולל Claude ו-GPT, הראה תוצאים לא אמינים. המודלים ייצרו רפרנסים שגויים בין 10.2% ל-98.4% מהפעמים.
תקציר מקורי באנגליתarXiv:2609.14988v1 Announce Type: new Abstract: Background. Large language models are increasingly used to help write biomedical text but may fabricate references to nonexistent work. How often large language models do so is not well characterized. Methods. We prompted 26 language models from eight developers (2023 to 2026) to supply a missing reference for each of 69 biomedical passages across ten domains. References were classified as verifiable (real paper with a resolving identifier), partial matches (real paper without a resolving identifier), fabricated (no matching indexed paper), or declined (the model refused to supply a reference). A reference was considered correct in every evaluated bibliographic field only when it was verifiable and its journal, year, and listed authors matche
קרא במקור המקורי