יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SpliTEE: ביצועי LLM מהירים ופרטיים על ידי קישור תשתיות עבודה מומשקות עם פרטיות דיפרנציאלית

SpliTEE: Fast and Private LLM Inference by Coupling GPU-Assisted Trusted Execution Environments with Differential Privacy
במאמר זה, המחברים מציגים את SpliTEE, ארכיטקטורה של ביצועי LLM מהירים ופרטיים. הם משתמשים בתשתיות עבודה מומשקות עם פרטיות דיפרנציאלית כדי להגן על פרטיות המשתמש. המחברים מציגים גם ניתוח עומק של תפקידי ה-LLM ומציעים דרך להגביר את הביצועים.
תקציר מקורי באנגליתarXiv:2609.15039v3 Announce Type: replace-cross Abstract: User prompts provided to large language models (LLMs) may contain private information. One way to protect them is to execute the LLM inside a trusted execution environment (TEE). However, this results in slow inference times as current TEEs are significantly slower than GPUs for LLM inference. To circumvent this, Tram\`er and Boneh (2019) proposed Slalom which splits neural network inference between a TEE and an untrusted GPU. They encrypt inputs to computations outsourced to the GPU. In this paper, we extend this split-inference architecture to LLM inference and instead protect intermediate inputs using differential privacy (DP). We first demonstrate that masking intermediate representations is necessary by showing an 80% accuracy
קרא במקור המקורי