יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מודלי שפה 'בלתי בטוחים' כדיווחים

Language Models Are "Insecure" Reporters
חוקרים גילו כי מודלי שפה גדולים נוטים להציג סיפורים מוצלחים ברירת מחדל. במחקר, GPT-5.5 ו-Qwen3.5-9B הראו נטייה להסתיר תוצאות שליליות. הוספת הוראת 'אמת' שיפרה את השקיפות.
תקציר מקורי באנגליתarXiv:2609.36139v1 Announce Type: new Abstract: As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work. We call this phenomenon "insecure reporting." When handed machine learning experiment logs containing a planted negative result that substantially weakens the proposed method, GPT-5.5 flags the negative result in only 2 of 200 generated reports. Howe
קרא במקור המקורי