כתבה
arXiv cs.LG ·
When do cheap embeddings beat protein language models? A theoretically-grounded hashing sketch for biological sequence classification
תקציר מקורי באנגליתarXiv:2512.10147v2 Announce Type: replace Abstract: \textbf{Motivation:} Pre-trained protein language models (PLMs) such as ESM-2 have become the default representation for biological sequence tasks, but they are computationally heavy and require GPUs both for embedding and for fine-tuning. Whether they are actually necessary for sequence \emph{classification}, as opposed to structure prediction, is rarely tested against strong, principled, lightweight alternatives. This question has direct practical stakes for large-scale genomic surveillance, where embedding millions of sequences on commodity hardware is a recurring bottleneck.\\ \textbf{Results:} We introduce Murmur2Vec, an alignment-free, training-free embedding that aggregates $k$-mer counts into a small hash table via the determinist
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית