כתבה
arXiv cs.AI ·
CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
CHERRY מציג חידוש באימון מודלי LLM: קיצור סדרתי של חוקי ייצוג עם רכיבי רפלקסיה. זה נעשה באמצעות סיפור-קפצה של סימן-אחסון, שמאפשר קיצור של 48 שכבות ל-6 חלקים ייחודיים, כאשר קיצור זה נשמר עד 227M.
תקציר מקורי באנגליתarXiv:2606.31796v2 Announce Type: replace-cross Abstract: Frontier language capability is usually bought with frontier compute; CHERRY shows a different trade. It is a sovereign Korean model family built on one principle: supervise the tokens that decide the answer, and let shared weights carry the rest. Under matched compute this exposes a sharp, reproducible dissociation---selected-token supervision preserves held-out discrimination yet collapses free generation, and a full-sequence anchor recovers only part of the gap. The same signal drives a heal-after-merge recurrent-representational-yield loop that collapses 48 layers to 6 unique blocks at near-dense parity (227M at loss 2.934 vs a 566M dense model at 2.926) and composes them by MoEE fusion (2.789)---a recurrent-compression directio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית