כתבה
arXiv cs.LG ·
SciTBERT: A family of chronologically consistent language models for scientific and technological language processing
תקציר מקורי באנגליתarXiv:2610.12207v1 Announce Type: cross Abstract: Pre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology. However, the applicability of these models for studying time-dependent or archival properties of science, technology, and their interface is limited due to lookahead and domain biases inherent to these pre-trained models. These limitations arise from training on corpora with unconstrained chronological and text source distributions. We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית