כתבה
arXiv cs.CL ·
BioBigBird: מודל תשומת לב רזה לעיבוד תלויות ארוכות טווח
BioBigBird: A Sparse Attention Model for Long-Range Dependency Processing in Biomedical Text
BioBigBird הוא מודל שפה בידירקציונלי שמותאם לעיבוד תלויות ארוכות טווח בטקסטים ביו-רפואיים. המודל משתמש במנגנון תשומת לב רזה ומאומן על מנת לייצר תוצאות תחרותיות.
תקציר מקורי באנגליתarXiv:2610.11430v1 Announce Type: new Abstract: While domain-specific Large Language Models (LLMs) have encoded vast biomedical knowledge, their limited context windows often hinder a deep understanding of nuanced relationships within and across texts. To address this limitation, we introduce BioBigBird, a bidirectional language model pre-trained on extensive biomedical literature and clinical data, specifically designed to handle long-range dependencies. BioBigBird leverages a sparse attention mechanism to process sequences up to 4096 tokens, and its training incorporates a multi-stage process to mitigate noise from the large-scale pre-training corpus. We further enhance its performance by employing a multi-task learning (MTL) framework that jointly optimizes for Named Entity Recognition
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית