יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

OmniTide: פיתוח יחד של אלגוריתמים ומערכות לשידור זרימה יעיל של LLM Omni-Modality

OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming
OmniTide מציע שידור זרימה יעיל של LLM Omni-Modality על גבי מכשירי קצה, עם זיכוי פרטיות וביצועי API נמוכים. הפיתוח כולל פיתוח יחד של אלגוריתם ומערכת, עם שיפורים במהירות ובאמינות.
תקציר מקורי באנגליתarXiv:2609.34653v2 Announce Type: replace Abstract: On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention methods fall short, either incurring prohibitive online estimation latency or destroying interleaved cross-modal context, while failing to resolve physical memory fragmentation. We present OmniTide, the first algorithm-system co-design tailored for efficient on-device streaming omni-modal inference. Driven by the observation of modality-aware structural sparsity, OmniTide adopts a unit-based abstraction with two components: (1) At the algorithm
קרא במקור המקורי