כתבה
arXiv cs.LG ·
דירוג היררכי במודלים גדולים של שפה
Hierarchical Grading in Large Language Models
חוקרים מציגים מודלים של שפה גדולים עם דירוג היררכי, המאפשרים שיפור בביצועים. המודלים משתמשים במרחב ייצוג עם דירוג ופועלים על נתונים באופן יעיל יותר.
תקציר מקורי באנגליתarXiv:2607.22757v1 Announce Type: new Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training objective. The construction extends the theory of graded neural networks and graded transformers to autoregressive language models while preserving expressive power, asymptotic computational complexity, and inference cost. The governing geometric picture is that of geometric invariant theory. The benefit of a grading is expressed by a Kempf--Ness functional on the grading torus; the grades that improve upon the uniform architecture form an open convex cone whose membership is decided by a Hilbert--Mum
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית