יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הפיצול של דגם המשתמש: דיסטילציה של תפיסה

User Model Extraction via Belief Self-Distillation
אנו מציגים פרוטוקול חדש של דיסטילציה של תפיסה, המאפשר פיצול של דגם המשתמש של LLM. הפרוטוקול, המכונה BSD, משתמש בטכניקה של דיסטילציה של תפיסה, כדי לפצל את התפיסה של ה-LLM על פי קונברסציות טבעיות.
תקציר מקורי באנגליתarXiv:2609.31603v1 Announce Type: new Abstract: Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stro
קרא במקור המקורי