כתבה
arXiv cs.LG ·
LGQ: Learnable Geometric Quantization לקיזום גאומטרי למודלי תמונה
LGQ: Learnable Geometric Quantization for Image Tokenization
LGQ: Learnable Geometric Quantization מציגה פתרון לקיזון גאומטרי למודלי תמונה. המאמר עוסק בפיתוח של LGQ, שהיא פתרון לקיזון גאומטרי שיכול להיות למודלי תמונה. LGQ משתמשת בטכנולוגיית Learnable Geometric Quantization כדי לקצור את התמונות לקודים קצרים. המאמר גם מציג תוצאות של LGQ בהשוואה לפתרונות אחרים.
תקציר מקורי באנגליתarXiv:2602.16086v4 Announce Type: replace-cross Abstract: Recent collapse-free quantizers such as FSQ achieve stable training by replacing the learnable codebook with an engineered geometry: a fixed scalar grid whose structure is dictated by the codebook size K. We show this trade-off is unnecessary. We introduce Learnable Geometric Quantization (LGQ), which retains a learnable codebook of codes and performs soft-to-hard assignment via temperature annealing, regularized by two cheap terms: a diversity term scaled by codebook size that penalizes concentrated batch-average usage is the primary driver of collapse resistance, complemented by a peakedness term that sharpens each token's soft-assignment toward one-hot; together they prevent codebook collapse without EMA, reset heuristics, or cod
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית