כתבה
arXiv cs.LG ·
LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling
תקציר מקורי באנגליתarXiv:2609.37675v1 Announce Type: new Abstract: Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and structure-based tasks, yet the potential of tokenization remains underexploited. Unlike human language, proteins preserve structure despite extensive sequence variation a property standard tokenization strategies fundamentally fail to capture. We introduce ZEST (Zoned Encoding of Sequence Traits), an evolution-informed vocabulary derived from conserved regions of multiple sequence alignments. ZEST allows embedding domain-level biological priors directly at the tokenization stage rather than learning them implicitly through scale. ZEST natively compresses sequences to an average token length of 4 r
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית