כתבה
arXiv cs.LG ·
הכללה של אורך לטרנספורמרים
Length Generalization for Transformers via Compression
חוקרים הצליחו לקבוע גבולות חדשים להכללה של אורך בטרנספורמרים, באמצעות שימוש במחרוזות מנוסחות. התגלית עשויה לשפר את הבנתנו את יכולות הלמידה של טרנספורמרים.
תקציר מקורי באנגליתarXiv:2609.08851v1 Announce Type: new Abstract: Recent advancements in transformer length generalization theory enable us to reliably predict when a transformer can learn to solve a task. In particular, the C-RASP hypothesis (a formalized version of the so-called RASP-l conjecture) posits that transformers length-generalize on a task if and only if a solution is expressible in the C-RASP language. While this hypothesis has strong empirical validation, theoretical problems arise from the fact that no computable length generalization bounds exist for C-RASP, alongside the discovery of seemingly contradictory experiments. To address these problems, we refine the C-RASP hypothesis utilizing the recently-proposed fragments C-RASP+ and C-RASP1. These fragments have computable length generalizati
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית