כתבה
arXiv cs.AI ·
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
תקציר מקורי באנגליתarXiv:2605.15913v5 Announce Type: replace-cross Abstract: Block attention, which processes the input as separate blocks that cannot attend to one another, offers significant potential to improve KV cache reuse in long-context scenarios such as Retrieval-Augmented Generation (RAG). However, its broader application is hindered by two key challenges: the difficulty of segmenting input text into meaningful, self-contained blocks, and the inefficiency of existing block fine-tuning methods that risk degrading performance. To address these, we first construct SemanticSeg, a large and diverse semantic segmentation dataset containing over 30k instances across 16 categories-including books, code, web text, and conversations with text lengths ranging from 2k to 32k. Using this dataset, we train a lig
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית