כתבה
arXiv cs.LG ·
BioBigBird: מודל של תשומת לב צפוף לעיבוד תלותיות ארוכות-טווח בביומדיקל טקסט
BioBigBird: A Sparse Attention Model for Long-Range Dependency Processing in Biomedical Text
מודל BioBigBird, שמשתמש בתשומת לב צפוף, מסוגל לעבד תלותיות ארוכות-טווח בביומדיקל טקסט. המודל נלמד על ספרות ביומדיקל ונתונים קליניים, ומציג תוצאות משופרות בשימוש במודלי רכיבה (MTL) לזיהוי גופים שמועברים וזיהוי יחסים.
תקציר מקורי באנגליתarXiv:2610.11430v1 Announce Type: cross Abstract: While domain-specific Large Language Models (LLMs) have encoded vast biomedical knowledge, their limited context windows often hinder a deep understanding of nuanced relationships within and across texts. To address this limitation, we introduce BioBigBird, a bidirectional language model pre-trained on extensive biomedical literature and clinical data, specifically designed to handle long-range dependencies. BioBigBird leverages a sparse attention mechanism to process sequences up to 4096 tokens, and its training incorporates a multi-stage process to mitigate noise from the large-scale pre-training corpus. We further enhance its performance by employing a multi-task learning (MTL) framework that jointly optimizes for Named Entity Recognitio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית