כתבה
arXiv cs.LG ·
LLM-based Source Code Compression via Thresholded Symbol Ranking
תקציר מקורי באנגליתarXiv:2607.24192v2 Announce Type: replace-cross Abstract: We study the problem of lossless compression of source code, motivated by the storage demands of large-scale software archives, such as Software Heritage (https://www.softwareheritage.org/). General-purpose compressors (e.g., zstd, bzip2) offer a good trade-off between compression ratio and speed, but fail to exploit all special regularities inherent in source code. Recent approaches leverage Large Language Models (LLMs) within Shannon's symbol-ranking framework, relying on a scheme in which the predicted rank can grow arbitrarily. While effective at reducing space, this setting incurs significant throughput degradation, and leaves open the question whether it is necessary to explicitly encode all ranks. In this work, we introduce L
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית