כתבה
arXiv cs.LG ·
Investigating Model Compression for Neural Machine Translation in the Biomedical Domain
תקציר מקורי באנגליתarXiv:2610.07032v1 Announce Type: cross Abstract: Large-scale pretrained transformer models have achieved state-of-the-art performance across diverse machine translation tasks, including multilingual settings. Knowledge distillation has emerged as a sustainable approach for model compression, transferring knowledge from large teacher models to smaller, more efficient student models. Similarly, quantization, which reduces the numerical precision of model weights and activations (e.g., from 32-bit to 8-bit representations) is widely used to accelerate inference, enabling models to run several times faster during deployment. However, both techniques face limitations when applied to specialized domain data, particularly under low-resource conditions. In knowledge distillation, the effectivenes
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית