כתבה
arXiv cs.AI ·
לימוד תצוגה סינתטית לסופרוויזיה ניגודית להצגת קוד
Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study
לימוד תצוגה סינתטית לסופרוויזיה ניגודית להצגת קוד. ניתן להשיג תוצאות טובות יותר בהשוואה לבסיסי טריינינג.
תקציר מקורי באנגליתarXiv:2609.03702v1 Announce Type: new Abstract: General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (setting-specific and costly to collect). We empirically study an alternative: contrastive pretraining of small encoders with synthetically generated natural-language descriptions emphasizing code functionality and intent, paired with code in a dual-encoder framework at training and discarded at inference. We benchmark this approach against pretraining-based baselines, generalist LLMs, and embedding-specific models on eight retrieval, classification, and generation tasks across C
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית