כתבה
MarkTechPost ·
JetBrains משחררת Mellum2.1: מודל פתוח 12B MoE לאגנטים קוד
JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents
JetBrains השיקה את Mellum2.1, מודל פתוח 12B MoE לאגנטים קוד. המודל, שפותח על ידי JetBrains, משמש לאגנטים קוד ומסוגל לעבוד באופן עצמאי. המודל נבנה על בסיס Mellum2, והוא כולל 28 שכבות ו-64 מומחים. המודל נטוע באמצעות RLHF, והוא כולל 131,072 תווים במסגרת הקשר. המודל נמצא ב-Hugging Face, והוא זמין תחת רישיון Apache 2.0.
תקציר מקורי באנגליתJetBrains has released Mellum2.1 , an open model built for coding agents and fast sub-agents. Mellum2.1 is a 12B mixture-of-experts thinking model from JetBrains that activates 2.5B parameters per token. It ships under Apache 2.0 on Hugging Face . The architecture is unchanged from Mellum2. The upgrade comes almost entirely from reinforcement learning (RL) in real software environments. The result is a small, self-hostable model that explores a repository, edits files, and checks its own changes. TL;DR Size: 12B total, 2.5B active (64 experts, 8 active), 131,072-token context Runs on: your own GPUs via vLLM or SGLang (speed tested on 1 NVIDIA H200). GGUF builds start at 7.0 GB for llama.cpp, Ollama and LM Studio. Performance: beats Mellum2 on 15 of 17 listed benchmarks; wins 5 of 17 agains
קרא במקור המקורי
marktechpost.com
פתח כתבה מקורית