כתבה
arXiv cs.AI ·
Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution
תקציר מקורי באנגליתarXiv:2610.07607v1 Announce Type: cross Abstract: Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide rich representations for this task, but task-agnostic zero-shot scores can be misaligned with a target assay, while supervised search in high-dimensional embedding spaces can make surrogate modeling and uncertainty estimation sample-inefficient. We propose the Linear Fitness Subspace (LFS) hypothesis: within mutation-induced residue-level representation changes, a compact, assay-specific set of directions makes fitness variation linearly accessible from few labeled variants. This is a local, supervision-recoverable statement rather than a claim that protein fitness landscapes or global PLM geom
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית