כתבה
arXiv cs.LG ·
התאמת ארכיטקטורות בזמן קווי למודלי טבלאות בהקשר
Adapting Linear-Time Architectures for Tabular In-Context Learning
במאמר זה, המחברים חוקרים את האפשרות לשימוש בארכיטקטורות בזמן קווי למודלי טבלאות בהקשר. הם מציעים פתרון לבעיית הגבלת המידע במודלים קלאסיים.
תקציר מקורי באנגליתarXiv:2609.36337v1 Announce Type: new Abstract: Tabular foundation models achieve strong performance by conditioning on labelled examples in context, but softmax attention limits their use on large datasets. Existing linear-time alternatives, however, are mostly causal, and their potential for tabular in-context learning (ICL) remains underexplored. To address this, we (1) revisit causal training setups, (2) compare linear sequence mixers, and (3) investigate their ICL generalisation beyond the pretraining context length. First, we show that the best training setup for causal models resembles next-token prediction. Then, perhaps surprisingly, the most promising linear sequence mixer is causal: DeltaNet outperforms even non-causal linear attention. However, it degrades beyond $2$-$4\times$
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית