יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מראה, מראה על הקיר: חזרה על הפרומפט במודלים קטנים

Mirror, Mirror on the Wall: Prompt Echoing in Small Instruct Language Models
חוקרים גילו תופעה של חזרה על הפרומפט במודלים קטנים כמו Llama, Qwen ו-SmolLM. התופעה נגרמת על ידי ראשי ההידוק של המודל, ולא על ידי הדליפה של תוכן המאמנים.
תקציר מקורי באנגליתarXiv:2609.15045v1 Announce Type: new Abstract: Prompt echoing is a recognized failure mode of instruct language models, in which a model instead of generating a response, mirrors the provided prompt, even though it did not receive a specific instruction to do so. Is this phenomenon a sign of the model leaking the content of its training dataset, or is it rather caused by a misaligned behavior of the internal induction/copying mechanisms? We investigate prompt echoing small language models from different families (Gemma, Llama, Qwen, SmolLM and OLMo) and show that echoing prompts are likely to have partial overlap with the training dataset but the phenomenon is primarily driven by the model's induction heads.
קרא במקור המקורי