יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

לימוד קומפרציות קומפרציה למטרות יעילות השגה במסגרת תרגום מכוני

Studying quantization trade-offs for efficient inference deployment in machine translation
במאמר זה, חוקרים חוקרים את קומפרציות הקומפרציה של EuroLLM למטרות יעילות השגה במסגרת תרגום מכוני. הם מדגימים כיצד קומפרציה יעילה יותר יכולה לשפר את ערך התרגום ולהפחית את זמן ההשגה.
תקציר מקורי באנגליתarXiv:2607.29397v3 Announce Type: replace Abstract: Deploying large language models in realistic server environments poses challenges, as the system needs to provide high-quality responses with low latency. Quantization is a common approach to reduce the memory footprint and improve inference efficiency, yet its impact on latency and throughput is rarely evaluated under controlled, orchestration-level workloads. In this work we study the quantization trade-offs of EuroLLM \citep{martins2025eurollm} across three model sizes ranging from 1.7B to 22B for efficient deployment on a single A100 or H100 GPU. We demonstrate that combining a document-chunking strategy with W4A8 or W8A8 quantization improves the latency-throughput Pareto-curve under a wide range of workloads. Furthermore, since stan
קרא במקור המקורי