כתבה
arXiv cs.CL ·
חוקי הקנה להזרקת ידע במודלי שפה גדולים
Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
חוקרים בודקים את היכולת של רשתות-על (hypernetworks) להזריק ידע למודלי שפה גדולים. הם מצאו שרשתות-על מסוגלות להזריק ידע בצורה מהימנה ובקנה מידה גדול. המחקר משתמש במודל ה-llama וב- LangChain.
תקציר מקורי באנגליתarXiv:2607.19604v1 Announce Type: new Abstract: Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge injection. Although hypernetworks are typically applied for test-time adaptation, we explore their use in train-time knowledge injection, where, given a large corpus of facts, we train a hypernetwork to generate a fixed LoRA adapter that, when inserted into the target model, enable the model to answer questions about those facts. In this work, we investigate whether hypernetworks can be used to perform train-time knowledge injection and how this ability varies with scale. The scaling behavior of hypernetworks remains largely unstudied. Our design decouples the hypern
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית