וידאו
YT AI Engineer ·
תורבוצ'רג' את זיכרון האג'נט עם TurboQuant
Turbocharge Your Agent's Retrieval with TurboQuant - Shashi Jagtap, Superagentic AI
▶ צפה כאן — בלי לצאת מהאתר
TurboQuant הוא שיטה לדחיסת וקטורים ל-3–4 ביטים, תוך שמירה על איכות. הוא פותח על ידי Google Research ויכול לשמש לשיפור זיכרון האג'נט.
תקציר מקורי באנגליתYour agent's memory has a hidden RAM bill. Every embedding and every token in the KV cache is stored at full 32-bit precision, four times heavier than search actually needs and it only grows as conversations and knowledge bases get bigger. Enter TurboQuant: a training-free compression method from Google Research (ICLR 2026) that squeezes each vector down to ~3–4 bits while keeping the one thing that matters which results rank closest to the query. In this talk we'll unpack how it works, why a tiny rerank step is the secret to keeping quality intact, and where it slots into your stack, from the model's KV cache to your RAG vector store. Then we'll prove it live: the same agent, the same answers, on an index ~5× smaller. You'll walk away with a practical, vendor-neutral way to make your agen
קרא במקור המקורי
youtube.com
פתח כתבה מקורית