יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ארכיטקטורה סימביוטית להרחבת אודיו

Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
החוקרים הציעו ארכיטקטורה סימביוטית להוספת יכולות אודיו למודלי שפה גדולים. הארכיטקטורה משתמשת במודול הזרקה שכותב וקטורים מותנים באודיו ישירות לזיכרון הקצר-טווח של המודל. השיטה מאפשרת למודל להתנהג כמודל שפה אודיו, מבלי לפגוע ביכולות המקוריות של המודל.
תקציר מקורי באנגליתarXiv:2609.30784v1 Announce Type: cross Abstract: This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme d
קרא במקור המקורי