יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SparLeak: פריצת פרטים מ-LLM בעזרת תשומת לב רזה

SparLeak: Privacy Leakage from Sparse Attention in LLM Inference on Shared GPUs
SparLeak מזהה פריצת פרטים ב-LLM בעזרת תשומת לב רזה, ומאפשר חשיפת מידע פרטי. המאמר מציג תצורה חדשה של תקלה בגישה לזיכרון, הנקראת SIMA, ומציע פתרון למניעת פריצת פרטים. SparLeak יכול לחשוף מידע פרטי, כולל תכונות של שאלות ותגובות של LLM.
תקציר מקורי באנגליתarXiv:2609.38830v1 Announce Type: new Abstract: Sparse attention is widely used to accelerate long-context inference in modern large language models (LLMs), but its input-dependent execution behavior introduces previously unexplored privacy risks. We identify a new GPU micro-architectural side channel, termed Sparsity-Induced Memory Access (SIMA), which arises from secret-dependent key-value cache access patterns induced by sparse attention. Based on this observation, we present SparLeak, a phase-aware side-channel attack that extracts SIMA traces during LLM inference and enables two practical privacy extractions: query attribute inference from prefill-phase traces and autoregressive response reconstruction from decoding-phase traces. By reconstructing approximate token-level sparsity prof
קרא במקור המקורי