יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

When do cheap embeddings beat protein language models? A theoretically-grounded hashing sketch for biological sequence classification

תקציר מקורי באנגליתarXiv:2512.10147v2 Announce Type: replace Abstract: \textbf{Motivation:} Pre-trained protein language models (PLMs) such as ESM-2 have become the default representation for biological sequence tasks, but they are computationally heavy and require GPUs both for embedding and for fine-tuning. Whether they are actually necessary for sequence \emph{classification}, as opposed to structure prediction, is rarely tested against strong, principled, lightweight alternatives. This question has direct practical stakes for large-scale genomic surveillance, where embedding millions of sequences on commodity hardware is a recurring bottleneck.\\ \textbf{Results:} We introduce Murmur2Vec, an alignment-free, training-free embedding that aggregates $k$-mer counts into a small hash table via the determinist
קרא במקור המקורי