יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה MarkTechPost ·

H Company משחררת NeoMME: משפחה של 260M ו-800M Single-Tower Multimodal Encoders

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
H Company השיקה NeoMME, משפחה של 260M ו-800M Single-Tower Multimodal Encoders שמשימתם לקצץ בפרמטרים ובחישוב.
תקציר מקורי באנגליתMost visual document retrievers in production today are hand-me-downs. ColPali and the models that followed it take a generative vision-language model and repurpose it as an encoder. The result still carries a separately pretrained vision tower and a causal decoder that never generates a token. That is parameter and compute overhead for a task that only needs representations. H Company has released NeoMME , a family of 260M and 800M bidirectional encoders that drops both components. One Transformer processes multilingual text tokens and raw 32×32 RGB image patches through the same layers, trained from random initialization. The retrieval fine-tune, NeoMME-Retriever, reaches 0.523 nDCG@10 on ViDoRe v3 at 260M parameters. Is it deployable? Yes. Every checkpoint ships under Apache 2.0 with da
קרא במקור המקורי