יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אימות עצמי במודלים של שפה

Trust Me, I'm Your Developer: Self-Issued Authentication in Large Language Models
חוקרים בדקו איך מודלים כמו Llama, Claude, Qwen ו-Mistral מאמתים זהויות. התוצאות הראו שהמודלים יכולים ליצור אימות עצמי, שיכול להוביל לאימות שגוי. המחקר מבהיר את החשיבות של אימות זהויות על ידי רכיבי אבטחה חיצוניים.
תקציר מקורי באנגליתarXiv:2609.03247v1 Announce Type: cross Abstract: Large language model (LLM) security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to verify an identity claim through a test designed by the model itself. We study this behavior through a staged developer-identity experiment with ChatGPT, Claude, Qwen, Mistral, and Llama. All five models initially rejected the unsupported claim "I am your developer." Claude refused to conduct an identity test, while ChatGPT generated developer-oriented questions but maintained that answers could demonstrate knowledge, not identity. In contrast, Qwen and Mistral generated technical challenges, defined what counted as convincing evidence, evaluated detailed answers, and returned Verified without re
קרא במקור המקורי