כתבה
arXiv cs.AI ·
Internalizer: פורטביליות של מפתח-לפרמטרים למודלי שפה גדולים
Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models
ה-Internalizer הוא פורטביליות של מפתח-לפרמטרים שמאפשרת למודלי שפה גדולים לשמור תכונות במשתנים שלהם. ה-Internalizer נבחן במודל DeepSeek v4 Flash.
תקציר מקורי באנגליתarXiv:2610.11715v1 Announce Type: new Abstract: Hypernetworks that map a context directly to a LoRA adapter let a large language model carry that context in its weights, but prior work has demonstrated them only on base models of up to 14 billion parameters. We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work. Most of its parameters live in a model-agnostic trunk with only thin entry and exit layers per base model, so it trains cheaply against small models before being ported to the large one. On unseen documents of up to 4096 tokens, the generated adapters reach 84.9% top-1 and 97.8
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית