יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

גילוי תוכן מזויף במודלים של LLM

When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination
חוקרים בדקו את יכולתו של מודל Google Gemini 3.0 Pro לגלות תוכן מזויף במסמכים. התוצאות הראו כי המודל מתקשה לגלות שגיאות כאשר הוא עובד עם קבצים גדולים. המודל גם הפיק תוצאות שגויות ובטוחות, כולל 'סנאי טלפתי' ו'מטוסר קוונטי'.
תקציר מקורי באנגליתarXiv:2609.09696v2 Announce Type: replace Abstract: Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical research, injecting 450 known contaminants of three types: typographical corruption, semantic reversal, and absurd out-of-context insertion. We then evaluate Google Gemini 3.0 Pro's ability to recover a 180-contaminant answer-key subset across 60 documents under three prompting regimes of increasing scale: single document, small batch, and large batch. Detection is unreliable even at small scale and collapses entirely at large scale: 50% recovery on single documents and 6
קרא במקור המקורי